Document Parse
The DocumentParse class aims to orchestrate the complete document processing pipeline, from extracting the raw content of the file to saving the processed data in a NoSQL database.
On this page
#Overview
The DocumentParse class aims to orchestrate the complete document processing pipeline, from extracting the raw content of the file to saving the processed data in a NoSQL database. It is responsible for bringing together several fundamental steps to turn digital files into structured data ready for use, automating the parsing flow with validations, cost calculations and integration with storage systems.
This component is fundamental in systems that need to interpret electronic documents (such as PDFs, TXT, etc.) and extract specific information according to a predefined schema. In practice, it can be used in machine learning pipelines, document conversion APIs, and applications that depend on automated data extraction.
With this class, the developer can submit a job containing a file and its settings, and get back the document's structured data with computational costs calculated, as well as having the result stored for future queries.
#Execution Flow
- Creating the class instance: Initialization with the job_id, JSON metadata, JSON schema, file content (in bytes), file extension and optional parsing settings.
- Calling the
run()method: This method runs the whole internal flow, and is responsible for coordinating the steps. - Loading and validating the data: The private method
_load_schema_and_config()deserializes the metadata and schema JSONs, and also merges the default and custom settings. - Extracting the file content:
_extract_file_content()uses theFileContentExtractorto obtain the raw text of the file according to its extension. - Parsing the extracted content:
_parse_content()invokes theContentParsingAgent, which maps the extracted text to the structure defined by the schema. - Calculating the costs:
_calculate_costs()determines the processing cost based on the language model used and the tokens consumed, converting the values to local currency. - Building the response:
_build_response()creates the payload that will be sent by the API and the payload that will be saved in the database. - Persisting the results:
_save()saves the processed data in the configured NoSQL database. - Return value:
run()returns the API-ready response containing the extracted and structured content.
#Class Methods Table
| Method | Description |
|---|---|
__init__ | Initializes the parsing job with the given data and settings |
run | Runs the complete document parsing pipeline |
_load_schema_and_config | Validates and loads the JSON metadata, schema and configuration |
_extract_file_content | Extracts the raw text of the file using the extractor |
_parse_content | Runs the agent that formats the content according to the schema |
_calculate_costs | Calculates the costs of the tokens used and converts the values |
_build_response | Builds the payloads for the API and the database |
_save | Persists the processed result in the NoSQL database |
#Environment Variables
No explicit environment variable is required directly by this class. The default configuration and database information are obtained through DocumentParseConfig, which may have its own ENV management (not shown here).
#Key Architecture Points and Insights
- Encapsulation and clear separation of responsibilities: The class organizes each step of the flow into private methods, making maintenance and testing easier.
- Use of composition: The class uses other classes for specific responsibilities, such as
FileContentExtractorfor extraction,ContentParsingAgentfor parsing andDocumentStorefor persistence, following the single responsibility principle. - Robust error handling: Uses HTTP exceptions to validate the JSONs provided, ensuring that the flow only continues with valid data.
- Dynamic cost calculation: Integrates with external services to calculate the cost based on the tokens processed and the current exchange rate conversion, suitable for pay-per-use systems.
- Configuration flexibility: Allows custom configuration via JSON that is merged with the default configuration to adapt the parsing behavior.
#Class and Methods Description
#DocumentParse Class
#Description
Class that manages the complete processing of documents to extract structured content. It receives the file content, metadata, schema and settings. It performs text extraction, parsing according to the schema, calculation of operational costs and storage of the processed result. It provides a response ready for use by APIs.
#Constructor Arguments
| Argument | Type | Description | Default Value |
|---|---|---|---|
job_id | str | Unique identifier of the processing job | - |
metadata | dict/str JSON | Metadata associated with the document (JSON string expected) | - |
schema | str | JSON string defining the schema of the expected data | - |
file_bytes | BytesIO | Binary content of the file to be processed | - |
file_extension | str | File extension (example: '.pdf', '.txt') | - |
config | Optional[str] | JSON string containing custom configuration for parsing | None |
#1. __init__
Description
Initializes the DocumentParse object with the essential information for processing the document, setting up parameters and loading the system's default settings.
Arguments
- job_id (str): Identifier of the processing job.
- metadata (dict/str): Metadata in JSON string format.
- schema (str): JSON schema for structuring the data.
- file_bytes (BytesIO): Bytes of the file.
- file_extension (str): File extension.
- config (Optional[str]): Custom configuration as a JSON string (optional).
Returns
- Returns no value.
Raises
- None.
Examples
document_parse = DocumentParse(
job_id="job123",
metadata='{"author": "John Doe"}',
schema='{"type": "object", "properties": {"title": {"type": "string"}}}',
file_bytes=BytesIO(b"conteudo do arquivo"),
file_extension=".txt",
config=None
)#2. run
Description
Runs the complete document processing pipeline, from data validation to extraction, parsing, calculations and storage, returning the final result ready for the API.
Arguments
- None.
Returns
- dict: Response formatted for the API containing the job_id and the processed content.
Raises
- HTTPException: in case of invalid JSONs in the data provided.
Examples
response = document_parse.run()
print(response)
# Exemplo de saída:
# {
# "job_id": "job123",
# "content": {...dados extraídos e formatados...}
# }#3. _load_schema_and_config
Description
Parses and validates the metadata, schema and configuration JSONs, merging the default configuration with the custom one when provided.
Arguments
- None.
Returns
- Returns no value.
Raises
- HTTPException: if the metadata, schema or config are invalid JSONs.
Examples
document_parse._load_schema_and_config()
# Inicializa self.metadata_data, self.schema_data e self.config_data com os valores carregados.#4. _extract_file_content
Description
Uses the FileContentExtractor class to extract the raw text content of the file based on its extension.
Arguments
- None.
Returns
- Returns no value.
Examples
document_parse._extract_file_content()
print(document_parse.result_extract)
# Exemplo de saída:
# {'response': 'Texto extraído do arquivo'}#5. _parse_content
Description
Runs the parsing agent ContentParsingAgent, which processes the extracted text and structures it according to the expected schema, preparing a formatted response.
Arguments
- None.
Returns
- Returns no value.
Examples
document_parse._parse_content()
print(document_parse.agent_response)
# Exemplo:
# {'content': {...dados estruturados...}, 'metadata': {...informações de processamento...}}#6. _calculate_costs
Description
Calculates the processing cost based on the input and output tokens used by the model, using the current exchange rate to convert to local currency (BRL).
Arguments
- None.
Returns
- Returns no value.
Examples
document_parse._calculate_costs()
print(document_parse.input_cost, document_parse.output_cost, document_parse.rate)
# Valores numéricos representando os custos#7. _build_response
Description
Builds the payloads that will be returned by the API and saved in the database, consolidating the extracted content and the cost and process information.
Arguments
- None.
Returns
- Returns no value.
Examples
document_parse._build_response()
print(document_parse.api_response)
# {'job_id': 'job123', 'content': {...dados extraídos...}}#8. _save
Description
Saves the processed data in the configured NoSQL database, using the DocumentStore class.
Arguments
- None.
Returns
- Returns no value.
Examples
document_parse._save()
# Dados persistidos no banco conforme configuração automática da classe#Use
if __name__ == "__main__":
# Exemplo de uso
job_id = "job_123"
metadata = """{"user_id": "user_456"}"""
schema = """
{
"summary": {
"type": "str",
"description": "Resumo do conteúdo do arquivo"
}
}
"""
config = """
{
"model_provider": "OpenAI",
"model_id": "gpt-4.1-mini",
"max_input_tokens": 1000000,
"debug_mode": true,
"instructions": "Extraia dados do texto",
"description": "Leia o texto e extraia as informações relevantes conforme o esquema definido. Retorne um JSON estruturado com os dados extraídos. Caso não encontre alguma informação, retorne null para aquele campo."
}
"""
with open("src\\content_parse\\module\\example.pdf", "rb") as f:
file_bytes = BytesIO(f.read())
parser = DocumentParse(
job_id=job_id,
metadata=metadata,
schema=schema,
config=config,
file_bytes=file_bytes,
file_extension="pdf"
)
response = parser.run()
print("\nResposta do parser:")
print(json.dumps(response, indent=2))
# python -m src.content_parse.module.document_parseSource: src/content_parse/module/document_parse.py