Retrieval Manager
On this page
#Overview
The RetrievalManager class is a tool for managing retrieved documents, aimed at text data search and analysis scenarios. It makes it easier to filter documents based on scores, to generate compact text contexts for use in language models and to extract file metadata in an organized way.
In practical applications, this class can be used to control large volumes of documents returned by information retrieval systems, ensuring that only the most relevant ones (with a score above a certain threshold or above the average) are considered. In addition, with the generation of formatted contexts, federating that content to LLMs or other text analyses becomes more efficient.
#Execution Flow
- Initialization: When creating a
RetrievalManagerinstance, you provide a list of documents, optionally a minimum score and whether you want score filtering to be done automatically. If automatic filtering is enabled, the internal list of documents will already be pre-filtered. - Filtering by Score: Documents can be filtered to keep only those whose score is greater than or equal to the desired minimum score, using
get_by_score. - Filtering by Mean: For refinement,
filter_by_mean_scoreallows filtering a set of chunks, keeping only those with a score greater than or equal to the set's mean. - Compact Context Generation:
generate_contextturns the documents into a line-by-line formatted string, showing score and text compactly to make it easier to use. - File Metadata Extraction: Through the
get_filesmethod, files associated with the documents are identified and organized to make it easier to access their identification, name, extension and maximum score obtained.
#Class Methods Table
| Method | Description |
|---|---|
__init__ | Initializes the instance with documents and settings |
get_by_score | Filters documents with a score greater than or equal to the minimum |
filter_by_mean_score | Filters chunks with a score greater than or equal to the mean |
generate_context | Generates a compact toon-style context string |
get_files | Extracts and organizes the files' metadata from the documents |
#Important Architecture Points and Insights
- Partial immutability: The class changes its internal list of documents only if automatic filtering is enabled, preserving the initial input when desired.
- Robust exception handling: All methods that do critical processing catch exceptions and raise
RuntimeErrorwith clear messages to make debugging easier. - Flexible document structure: The class assumes that documents have a standard format with required fields such as 'text' and 'score', and optional 'metadata' fields with file information — which allows adaptation to different data sources.
- No external dependencies: It does not use other helper classes, which makes it isolated and simple to use, only manipulating lists and dictionaries.
- Distinct filters: With distinct methods for fixed-threshold filtering and mean filtering, management is more flexible for different relevance criteria.
#Class and Methods Description
#RetrievalManager Class
#Description
Manages a collection of text documents with associated scores, allowing multiple important operations:
- Filter documents based on scores (fixed or relative to the group's mean).
- Create a compact text representation for integration with language models.
- Extract refined metadata to associate the documents with files.
Serves cases that involve advanced retrieval and handling of texts organized by relevance.
#Constructor Arguments
| Argument | Type | Description | Default Value |
|---|---|---|---|
docs | List[Dict[str, Any]] | List of documents with the fields 'text', 'score' and optional 'metadata' | - |
score_min | float | Minimum score value for the initial filtering of documents | 0.0 |
filter_by_score | bool | Defines whether documents will be filtered by score at initialization | False |
#Methods
#1. __init__
#Description
Builds the RetrievalManager object with a list of documents and optional settings for initial filtering by minimum score.
#Arguments
- docs (List[Dict[str, Any]]): documents to manage
- score_min (float): minimum score for the initial filter
- filter_by_score (bool): runs the automatic filter if True
#Returns
- Does not return a value.
#Raises
- Not applicable.
#Examples
rm = RetrievalManager(docs=documents, score_min=0.3, filter_by_score=True)#2. get_by_score
#Description
Filters the list of documents to keep only those whose score is greater than or equal to the given minimum score.
#Arguments
- docs (List[Dict[str, Any]]): Optional list of documents to filter. If omitted, uses the internal list.
- score_min (float): Minimum score for filtering.
#Returns
- List[Dict[str, Any]]: filtered list of documents.
#Raises
- RuntimeError: when the score filter fails.
#Examples
filtered_docs = rm.get_by_score(score_min=0.35)#3. filter_by_mean_score
#Description
Receives a list of chunks (documents), calculates the mean of the scores and filters to keep only the chunks with a score greater than or equal to that mean.
#Arguments
- chunks (list[dict[str, Any]]): list of chunks with the 'score' field.
#Returns
- list[dict[str, Any]]: filtered list, or empty if the input is empty.
#Raises
- RuntimeError: if an error occurs in the calculation or filtering.
#Examples
chunks = [
{"text": "A", "score": 0.5},
{"text": "B", "score": 0.3},
{"text": "C", "score": 0.7},
]
filtered = rm.filter_by_mean_score(chunks)
# filtered conterá chunks com score >= 0.5#4. generate_context
#Description
Generates a compact string in the format "Score: X | Content: text" for each document, replacing line breaks in the text with spaces for better formatting.
#Arguments
- docs (List[Dict[str, Any]]): documents to be converted to context. Uses the class's internal ones if omitted.
#Returns
- str: concatenated string with the documents formatted line by line.
#Raises
- RuntimeError: on failure in processing the string.
#Examples
context_string = rm.generate_context()
print(context_string)
# Exemplo saída:
# Score: 0.38 | Content: xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
# Score: 0.36 | Content: xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx#5. get_files
#Description
Extracts standardized information about files present in the documents' metadata, returning a list of files with id, name, extension and highest associated score.
#Arguments
- docs (List[Dict[str, Any]]): list of documents to extract metadata from. Uses the internal list if omitted.
#Returns
- List[Dict[str, Any]]: list of dictionaries representing files filtered by id and ordered according to the last processing.
#Raises
- RuntimeError: in case of an error extracting metadata.
#Examples
files_info = rm.get_files()
print(files_info)
# [
# {"id": "cucinare", "name": "LESSICO per CUCINARE.pdf", "ext": "pdf", "score": 0.38},
# {"id": "tenerezza", "name": "TENEREZZA.pdf", "ext": "pdf", "score": 0.36}
# ]#Usage
if __name__ == "__main__":
import json
documents = [
{
"id": "79258322-c06b-4e50-9a69-c8caa1136b3f",
"text": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx",
"metadata": {
"Client": "1234",
"client_id": "0011",
"collection_id": "collection_01",
"collection_name": "BetterAI",
"created_at": "2026-03-25 18:58:43",
"file_extension": "pdf",
"file_id": "cucinare",
"file_name": "LESSICO per CUCINARE.pdf",
"user_id": "11"
},
"score": 0.382682741
},
{
"id": "f75b0f7c-36ab-48d4-8da0-ec21b9ce688a",
"text": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx",
"metadata": {
"Client": "1234",
"client_id": "0011",
"collection_id": "collection_01",
"collection_name": "BetterAI",
"created_at": "2026-03-25 18:58:43",
"file_extension": "pdf",
"file_id": "tenerezza",
"file_name": "TENEREZZA.pdf",
"user_id": "11"
},
"score": 0.359430343
},
{
"id": "f75b0f7c-36ab-48d4-8da0-ec21b9ce688a",
"text": "xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx",
"metadata": {
"Client": "1234",
"client_id": "0011",
"collection_id": "collection_01",
"collection_name": "BetterAI",
"created_at": "2026-03-25 18:58:43",
"file_extension": "pdf",
"file_id": "tenerezza",
"file_name": "TENEREZZA.pdf",
"user_id": "11"
},
"score": 0.329430343
},
]
meneger = RetrievalManager(
docs=documents,
score_min=0.36,
#filter_by_score=True
)
print(json.dumps(meneger.get_files(), indent=2))
# python -m src.vector_store.pinecone.utils.retrieval_managerThis documentation details the practical operation of the RetrievalManager class, serving as a guide for developers who need to handle documents with relevance metrics and associated metadata.
Source: src/vector_store/pinecone/utils/retrieval_manager.py