Image Generation

Image Generator Service

Service that orchestrates image generation and editing with Gemini: builds the multimodal prompt, configures the generation, calls the model and parses the responses.

On this page

    Source: src/image_generation/image_generator_service.py. This page's original notebook was empty; the content was written from the code.

    #Overview

    The ImageGeneratorService class covers four steps, in this order: building the parts (text and images), configuring the generation, calling the model and parsing the responses (images, text and usage metrics).

    #Constructor

    ParameterTypeDescription
    clientgenai.Client (optional)Gemini client already initialized. If None, a new one is created with GeminiClient().get_client().
    content_configdict (optional)Overrides values of DEFAULT_CONTENT_CONFIG (model, temperature, top_p, aspect_ratio, number_of_images, etc.).

    The final configuration is the merge of DEFAULT_CONTENT_CONFIG with content_config: missing keys inherit the default value. Any failure during initialization raises RuntimeError.

    #Methods

    #build_parts

    Builds the list of types.Part sent to Gemini. The text is made of BASE_PROMPT, followed by [TASK] with the user_prompt and, if instructions is set, by [USER_CONTEXT]. Each reference image in images is appended as a bytes part.

    An empty user_prompt, or one that is not a str, raises an error (ValueError, re-raised as RuntimeError). The method also stores the prompt, the instructions and the image count in self.text_input, for later use in parse_responses.

    #generate_config

    Returns a types.GenerateContentConfig with temperature, top_p and max_output_tokens from the configuration, response_modalities equal to ["IMAGE", "TEXT"] and ImageConfig with the aspect_ratio.

    #call_model

    Calls client.models.generate_content once for each requested image, with the same prompt, and returns the list of responses. The number of calls is content_config["number_of_images"].

    Cost and limits. The current Gemini SDK does not configure the number of images per call; this is why generating several images is done with several calls. Avoid high values of number_of_images to avoid exceeding rate limits or raising the cost.

    #parse_responses

    Walks through the responses and aggregates the result into a single dictionary:

    KeyContent
    text_inputPrompt, instructions and image count (only present after build_parts).
    text_responsesList with the texts returned by the model.
    imagesList of dictionaries with mime_type and data (bytes).
    generate_configThe final configuration used for the generation.
    usage_metadataOne entry per response, with prompt_tokens, output_tokens and total_tokens, or None when the response carries no metrics.

    #Usage example

    from src.image_generation.image_generator_service import ImageGeneratorService
    
    # Create the service with a partial configuration
    service = ImageGeneratorService(content_config={"number_of_images": 2})
    
    # 1. Build the parts, 2. configure, 3. call the model, 4. parse
    parts = service.build_parts(user_prompt="Um grifo renascentista")
    config = service.generate_config()
    responses = service.call_model(parts, config)
    result = service.parse_responses(responses)
    
    print(len(result["images"]))

    #Notes

    Input image type. In build_parts, all reference images are sent with mime_type image/jpeg, even when the original file is a PNG.

    See also Config, Gemini Client and Module.

    Source: src/image_generation/image_generator_service.py

    Esc
    ↑↓navigate Enteropen