Parse Content

Document Parse

Receives a file and a JSON schema to extract and structure its content via LLM.

On this page

    #Document Parse - /parse-content/document-parse

    POST/parse-content/document-parse

    Receives a file and a JSON schema to extract and structure its content via LLM. Use it when you need to turn unstructured documents (PDFs, Word, text) into organized data ready for consumption.

    #Parameters

    ParameterTypeDescriptionExample
    metadatastring (JSON)Auxiliary metadata of the document, free for the client to use.{"client": "Acme", "doc_type": "contract"}
    document_schemastring (JSON)Schema that defines the fields to extract. Each key takes type and description.{"summary": {"type": "str", "description": "Resumo do documento"}}
    filefile (upload)File to be processed. Accepted formats: txt, md, pdf, docx. Limit: 50MB.contrato.pdf
    configstring (JSON) — optionalParser settings: LLM model, custom instructions and debug mode.{"model_provider": "OpenAI", "model_id": "gpt-4.1-mini", "debug_mode": false}

    Note: the job_id is generated automatically by the server and returned in the response.

    Schema formats: https://github.com/enzoschitini/better-ai/blob/production/doc/Modules/Content%20Parse/Json%20To%20Pydantic.md

    #Config (explanation of each parameter)

    The config field is optional. When not sent, the API uses the parser's internal default_config.

    FieldTypeWhat it doesDefaultWhen to adjust
    model_providerstringDefines the language model provider. Currently the main flows use OpenAI and Groq."OpenAI"Change it when you want to switch provider because of cost, latency or availability.
    model_idstringIdentifies the model used for extraction (gpt-4.1-mini, for example)."gpt-4.1-mini"Adjust it when you need higher quality, lower cost or a different context window.
    max_input_tokensintegerLimit of input tokens allowed to validate the size of the context before execution.1000000Lower it to impose a cost limit or raise it in scenarios with very long documents.
    debug_modebooleanTurns on logs and debugging information of the parsing agent.falseUse true for troubleshooting and go back to false in production.
    instructionsstringOperational instructions on how to extract the data (style, rules, focus)."Extraia dados do texto"Customize it to guide the model toward formats and criteria specific to your domain.
    descriptionstringRole/expected behavior of the agent, reinforcing the response format and the fallback with null."Leia o texto e extraia as informações relevantes conforme o esquema definido..."Adjust it for business scenarios with their own filling and validation policies.

    Complete config example:

    {
      "model_provider": "OpenAI",
      "model_id": "gpt-4.1-mini",
      "max_input_tokens": 1000000,
      "debug_mode": false,
      "instructions": "Extraia dados do texto",
      "description": "Leia o texto e extraia as informações relevantes conforme o esquema definido. Retorne um JSON estruturado com os dados extraídos. Caso não encontre alguma informação, retorne null para aquele campo."
    }

    #Request

    curl --location 'http://localhost:8000/parse-content/document-parse' \
    --header 'X-API-Key: ******' \
    --form 'metadata="{\"value1\": \"value3\"}"' \
    --form 'document_schema="{
      \"summary\": {
        \"type\": \"str\",
        \"description\": \"Resumo do conteúdo do arquivo\"
      }
    }"' \
    --form 'config="{
      \"model_provider\": \"OpenAI\",
      \"model_id\": \"gpt-4.1-mini\",
      \"debug_mode\": true,
      \"instructions\": \"Extraia dados do texto\",
      \"description\": \"Leia o texto e extraia as informações relevantes conforme o esquema definido. Retorne um JSON estruturado com os dados extraídos. Caso não encontre alguma informação, retorne null para aquele campo.\"
    }"' \
    --form 'file=@"/path/to/file"'

    #Response

    {
        "job_id": "job_1788198742619260200fKBY",
        "status": "success",
        "status_code": 200,
        "content": {
            "summary": "..."
        },
        "time": {
            "start": "2026-08-31 14:52:22",
            "end": "2026-08-31 14:52:30",
            "duration_seconds": 8.0
        }
    }

    Source: src/web_services_network/routes/parse_content.py

    Esc
    ↑↓navigate Enteropen