Image Generator Service
Service that orchestrates image generation and editing with Gemini: builds the multimodal prompt, configures the generation, calls the model and parses the responses.
On this page
Source: src/image_generation/image_generator_service.py. This page's original notebook was empty; the content was written from the code.
#Overview
The ImageGeneratorService class covers four steps, in this order: building the parts (text and images), configuring the generation, calling the model and parsing the responses (images, text and usage metrics).
#Constructor
| Parameter | Type | Description |
|---|---|---|
client | genai.Client (optional) | Gemini client already initialized. If None, a new one is created with GeminiClient().get_client(). |
content_config | dict (optional) | Overrides values of DEFAULT_CONTENT_CONFIG (model, temperature, top_p, aspect_ratio, number_of_images, etc.). |
The final configuration is the merge of DEFAULT_CONTENT_CONFIG with content_config: missing keys inherit the default value. Any failure during initialization raises RuntimeError.
#Methods
#build_parts
Builds the list of types.Part sent to Gemini. The text is made of BASE_PROMPT, followed by [TASK] with the user_prompt and, if instructions is set, by [USER_CONTEXT]. Each reference image in images is appended as a bytes part.
An empty user_prompt, or one that is not a str, raises an error (ValueError, re-raised as RuntimeError). The method also stores the prompt, the instructions and the image count in self.text_input, for later use in parse_responses.
#generate_config
Returns a types.GenerateContentConfig with temperature, top_p and max_output_tokens from the configuration, response_modalities equal to ["IMAGE", "TEXT"] and ImageConfig with the aspect_ratio.
#call_model
Calls client.models.generate_content once for each requested image, with the same prompt, and returns the list of responses. The number of calls is content_config["number_of_images"].
Cost and limits. The current Gemini SDK does not configure the number of images per call; this is why generating several images is done with several calls. Avoid high values of number_of_images to avoid exceeding rate limits or raising the cost.
#parse_responses
Walks through the responses and aggregates the result into a single dictionary:
| Key | Content |
|---|---|
text_input | Prompt, instructions and image count (only present after build_parts). |
text_responses | List with the texts returned by the model. |
images | List of dictionaries with mime_type and data (bytes). |
generate_config | The final configuration used for the generation. |
usage_metadata | One entry per response, with prompt_tokens, output_tokens and total_tokens, or None when the response carries no metrics. |
#Usage example
from src.image_generation.image_generator_service import ImageGeneratorService
# Create the service with a partial configuration
service = ImageGeneratorService(content_config={"number_of_images": 2})
# 1. Build the parts, 2. configure, 3. call the model, 4. parse
parts = service.build_parts(user_prompt="Um grifo renascentista")
config = service.generate_config()
responses = service.call_model(parts, config)
result = service.parse_responses(responses)
print(len(result["images"]))#Notes
Input image type. In build_parts, all reference images are sent with mime_type image/jpeg, even when the original file is a PNG.
See also Config, Gemini Client and Module.
Source: src/image_generation/image_generator_service.py