Image Generation
Image Generation
Generates images from a text prompt, with support for style instructions, model settings and reference images.
On this page
#Image Generation - /davinci/image-generation
POST/davinci/image-generation
Generates images from a text prompt, with support for style instructions, model settings and reference images. Use it when you need to create or transform images programmatically via a multimodal LLM.
#Parameters
| Parameter | Type | Description | Example |
|---|---|---|---|
user_input | string | Main prompt describing the image to be generated. | "Uma pizzaria napolitana à noite" |
instructions | string — optional | Additional style instructions or creative guidelines for the generation. | "Estilo de animação 3D como as da Disney" |
config | string (JSON) — optional | Model settings: model, temperature, top_p, max_output_tokens, aspect_ratio, number_of_images. | {"model": "gemini-2.5-flash-image", "aspect_ratio": "9:16"} |
files | file[] — optional | Reference images to guide the generation. Accepts multiple uploads. | reference.png |
#Config (explanation of each parameter)
The config field is optional. When not sent, the API uses the generation module's internal DEFAULT_CONTENT_CONFIG. The fields sent override only the corresponding values.
| Field | Type | What it does | Default | When to adjust |
|---|---|---|---|---|
model | string | Multimodal model used for generation (gemini-2.5-flash-image, gemini-3-pro-image-preview). | "gemini-2.5-flash-image" | Switch to the pro model when you need higher quality or more output tokens. |
temperature | float (0.0–2.0) | Controls the creativity of the generation. Low values = more predictable results. | 0.75 | Lower it for greater fidelity to the prompt, raise it for more creative variation. |
top_p | float (0.0–1.0) | Nucleus sampling: limits the set of tokens considered. | 0.85 | Lower it for more conservative and consistent outputs. |
max_output_tokens | integer | Output token limit. gemini-2.5-flash-image: 1024 or 2048. gemini-3-pro-image-preview: 4096 or 8192. | 1024 | Raise it when using pro models or when the text response is truncated. |
aspect_ratio | string | Aspect ratio of the generated image. See accepted values below. | "1:1" | Adjust it according to the image's destination (feed, stories, banner). |
number_of_images | integer (≥ 1) | Number of images generated. Each image is an independent call to the model. | 1 | Raise it to get variations, remembering that cost and latency grow proportionally. |
#Accepted values for aspect_ratio
| Value | Format | Typical use |
|---|---|---|
1:1 | Square | Feed post, avatar, thumbnail |
3:4 | Portrait | Portraits, product catalog |
4:3 | Landscape | Classic images, presentations |
9:16 | Full portrait | Stories, Reels, TikTok |
16:9 | Widescreen | Banners, covers, video/hero |
#Request
curl --location 'http://localhost:8000/davinci/image-generation' \
--form 'user_input="Crie a imagem de uma pizzaria napolitana"' \
--form 'instructions="O estilo deve ser uma animação 3D como as da Disney"' \
--form 'config="{
\"model\": \"gemini-2.5-flash-image\",
\"temperature\": 0.75,
\"top_p\": 0.85,
\"max_output_tokens\": 1024,
\"aspect_ratio\": \"9:16\",
\"number_of_images\": 2
}"' \
--form 'files=@"/path/to/file"'#Response
{
"job_id": "job_1788272755499472800EkNu",
"status": "success",
"status_code": 200,
"result": {
"text_response": "Com certeza! Aqui está uma pizzaria napolitana em estilo de animação 3D da Disney:\n\n \n",
"images": [
"https://hsenyunovbrmjejxqvjn.supabase.co/storage/v1/object/public/images/image_generations/img_1788272774157403200z8UW",
"https://hsenyunovbrmjejxqvjn.supabase.co/storage/v1/object/public/images/image_generations/img_17882727741574032007MBY"
]
},
"time": {
"start": "2026-09-01 11:25:55",
"end": "2026-09-01 11:26:19",
"duration_seconds": 24
}
}Source: src/web_services_network/routes/davinci.py