Deep Research / Tavily Research

Context Builder

Documentation of the TavilyContextBuilder and TavilyResearchRunner classes and of the endpoint that turn a web search into markdown context for language models.

On this page

    #Overview

    The context_builder module turns a web search done through Tavily into a markdown context ready to be used by language models, along with the list of URLs analyzed. It is made up of two classes:

    • TavilyContextBuilder: runs the pipeline (search, score filtering, conversion to markdown and URL extraction).
    • TavilyResearchRunner: builds the search configuration, applying default values, and triggers the TavilyContextBuilder.

    The search itself is done by a "researcher" object, normally TavilyDeepResearch. The pipeline is exposed by the API at the POST /deep-research/context-builder endpoint.

    #Execution Flow

    1. Search configuration: TavilyResearchRunner.run() calls build_search_config(), which validates query and max_results, applies the default values (search_depth, topic and include_answer) and incorporates extra parameters.
    2. Pipeline: the runner hands the configuration to TavilyContextBuilder.build_context().
    3. Search: search() calls researcher.start_search() with the configuration.
    4. Filtering: filter_results() keeps only the results with a score greater than or equal to min_score, removes the disposable global keys and the raw_content of each result.
    5. Markdown: to_markdown() converts the filtered response into markdown text.
    6. URLs: get_web_sites() gathers the URLs of the results kept.
    7. Return: build_context() returns a dictionary with the keys markdown and urls.
    8. Errors: at each stage, an exception is recorded with ApplicationTracing (ERROR level) and re-raised as a RuntimeError with a descriptive message.

    #Classes Methods Table

    ClassMethodDescription
    TavilyContextBuilder__init__Defines the researcher, the minimum score and the keys to remove.
    TavilyContextBuildersearchRuns the search with the configuration received.
    TavilyContextBuilderfilter_resultsFilters the results by score and removes unnecessary fields.
    TavilyContextBuilderto_markdownConverts the filtered results into markdown.
    TavilyContextBuilderget_web_sitesReturns the list of the results' URLs.
    TavilyContextBuilderbuild_contextRuns the complete pipeline.
    TavilyResearchRunner__init__Defines the builder and the search's default values.
    TavilyResearchRunnerbuild_search_configBuilds the search configuration dictionary.
    TavilyResearchRunnerrunBuilds the configuration and runs the pipeline.

    #Environment Variables

    • TAVILY_API_KEY: Tavily API key, read by TavilyDeepResearch when it is instantiated.

    #Important Architecture Points and Insights

    • Separation of responsibilities: TavilyResearchRunner takes care of configuration and defaults; TavilyContextBuilder takes care of processing; the researcher takes care of access to Tavily.
    • Dependency injection: the builder receives the researcher and the runner receives the builder, which makes it easy to swap any piece.
    • Relevance filtering: only the results with a score greater than or equal to min_score (default 0.5) enter the context.
    • Lean context: the response's response_time, follow_up_questions, images and request_id keys and each result's raw_content are removed before conversion.
    • Uniform error handling: every failure is recorded with ApplicationTracing (flag "Deep Research") and re-raised as a RuntimeError.
    • Informational logs disabled: the tracer is created with show_info_logs=False, and the module's tracer.INFO calls are commented out.

    #Generated Markdown Format

    The to_markdown() method produces a text with this structure, repeating the source block for each result:

    # Consulta
    **Pergunta:**
    {query}
    
    ---
    
    ## Resposta resumida
    {answer}
    
    ---
    
    ## Fontes analisadas
    
    ### 1. {title}
    - **URL:** {url}
    - **Score de relevância:** {score}
    
    **Conteúdo:**
    {content}
    
    ---

    The score is rounded to 5 decimal places. If a result has no title, "Sem título" is used.

    #API Endpoint

    POST/deep-research/context-builder

    The endpoint builds a TavilyDeepResearch, a TavilyContextBuilder (with the min_score received) and a TavilyResearchRunner, and returns the result of run() in the result field of the standard response. It requires the BetterAI API key in the X-API-Key header (validation is skipped when LOCAL=true).

    #Request body

    FieldTypeDescriptionDefault Value
    querystrMain query or subject of the research."What are the latest AI trends in the healthcare industry?"
    search_depth"basic" | "advanced"Depth of the research: basic for a quick overview, advanced for a more comprehensive search."advanced"
    max_resultsint (1 to 100)Maximum number of results to retrieve and include in the context.35
    topicstrCategory or domain of the research. Examples: general, news, finance."general"
    include_answerboolWhether to include an answer generated from the results.True
    min_scorefloat (0.0 to 1.0)Minimum relevance score; results below it are discarded.0.5

    #Request example

    curl -X POST "http://127.0.0.1:8000/deep-research/context-builder" \
      -H "X-API-Key: SUA_CHAVE" \
      -H "Content-Type: application/json" \
      -d '{
        "query": "Quais as principais tendências de IA em 2026?",
        "search_depth": "advanced",
        "max_results": 10,
        "topic": "general",
        "include_answer": true,
        "min_score": 0.5
      }'

    #Success response

    The response follows the API's standardized format, with job_id, status, result and time metrics. In result come the markdown and the URLs:

    {
      "job_id": "job_...",
      "status": "success",
      "status_code": 200,
      "result": {
        "markdown": "# Consulta\n...",
        "urls": ["https://..."]
      },
      "time": {
        "start": "2026-01-01 10:00:00",
        "end": "2026-01-01 10:00:05",
        "duration_seconds": 5.0
      }
    }

    In case of failure, the response carries status equal to error, status_code 500 and an error object with type and message. The API's full parameters and examples are in Web Service Network - API.

    #Classes and Methods Description

    #TavilyContextBuilder Class

    #Description

    Runs the pipeline that turns a web search into markdown context: search, score filtering, conversion to markdown and URL extraction.

    #Constructor Arguments

    ArgumentTypeDescriptionDefault Value
    researcherobjectObject with the start_search method, normally a TavilyDeepResearch.—
    min_scorefloatMinimum score to keep a result.0.5
    remove_keysList[str] | NoneKeys removed from the response's top level.["response_time", "follow_up_questions", "images", "request_id"]

    #Methods

    #Description

    Runs the search by calling researcher.start_search() with the configuration received.

    #Arguments

    • search_config (Dict[str, Any]): search configuration (for example, query, max_results, search_depth).

    #Returns

    • Dict[str, Any]: raw search response.

    #Raises

    • RuntimeError: if the search fails ("Error during web search").

    #Examples

    raw = builder.search({"query": "tendências de IA em 2026", "max_results": 5})

    #2. filter_results

    #Description

    Keeps only the results whose score is greater than or equal to min_score, removes the global keys defined in remove_keys and removes each result's raw_content.

    #Arguments

    • research_results (Dict[str, Any]): raw search response.

    #Returns

    • Dict[str, Any]: filtered response.

    #Raises

    • RuntimeError: if the filtering fails ("Error filtering web search results").

    #Examples

    filtered = builder.filter_results(raw)

    #3. to_markdown

    #Description

    Converts the filtered response into markdown text, with the query, the summarized answer and the sources analyzed (title, URL, relevance score and content).

    #Arguments

    • data (Dict[str, Any]): filtered response.

    #Returns

    • str: markdown text.

    #Raises

    • RuntimeError: if the conversion fails ("Error converting web search results to markdown").

    #Examples

    markdown = builder.to_markdown(filtered)

    #4. get_web_sites

    #Description

    Returns the list with the URL of each result in the filtered response.

    #Arguments

    • filtered_result (dict): filtered response, with the results key.

    #Returns

    • list: URLs of the results.

    #Examples

    urls = builder.get_web_sites(filtered)

    #5. build_context

    #Description

    Runs the complete pipeline: search, filter_results, to_markdown and get_web_sites.

    #Arguments

    • search_config (Dict[str, Any]): search configuration.

    #Returns

    • dict: dictionary with the keys markdown (markdown text) and urls (list of URLs analyzed).

    #Raises

    • RuntimeError: if any stage fails ("Error building research context").

    #Examples

    context = builder.build_context({"query": "tendências de IA em 2026", "max_results": 5})
    print(context["markdown"])
    print(context["urls"])

    #TavilyResearchRunner Class

    #Description

    Builds the search configuration, applying default values, and triggers the TavilyContextBuilder. It is the simplest entry point for generating a context.

    #Constructor Arguments

    ArgumentTypeDescriptionDefault Value
    builderobjectObject with the build_context method, normally a TavilyContextBuilder.—
    default_search_depthstrDepth used when search_depth is not provided."advanced"
    default_topicstrTopic used when topic is not provided."general"
    default_include_answerboolValue used when include_answer is not provided.True

    #Methods

    #1. build_search_config

    #Description

    Builds the search configuration dictionary, validating the required fields, applying the default values and incorporating extra_params.

    #Arguments

    • query (str): query. Required.
    • max_results (int): maximum number of results. Required.
    • search_depth (Optional[str]): search depth. If omitted, uses the runner's default.
    • topic (Optional[str]): search topic. If omitted, uses the runner's default.
    • include_answer (Optional[bool]): whether to include the summarized answer. If omitted, uses the runner's default.
    • extra_params (Optional[Dict[str, Any]]): additional parameters incorporated into the configuration; they can override the fields above.

    All arguments are named (keyword-only).

    #Returns

    • Dict[str, Any]: search configuration.

    #Raises

    • ValueError: if query is empty or if max_results is None.

    #Examples

    config = runner.build_search_config(
        query="tendências de IA em 2026",
        max_results=5,
        extra_params={"include_domains": ["arxiv.org"]},
    )

    #2. run

    #Description

    Builds the configuration with build_search_config and runs the pipeline with builder.build_context.

    #Arguments

    • The same as build_search_config, all named.

    #Returns

    • dict: dictionary with the keys markdown and urls, as in build_context.

    #Raises

    • ValueError: if query or max_results are not provided.
    • RuntimeError: if any stage of the pipeline fails.

    #Examples

    researcher = TavilyDeepResearch()
    
    builder = TavilyContextBuilder(
        researcher=researcher,
        min_score=0.5
    )
    
    runner = TavilyResearchRunner(builder)
    
    context = runner.run(
        query="Quais as principais tendências de IA em 2026?",
        search_depth="advanced",
        max_results=2,
        topic="general",
        include_answer=True
    )
    
    print(context["markdown"])
    # python -m src.deep_research.tavily_research.context_builder

    Source: src/deep_research/tavily_research/context_builder.py

    Esc
    ↑↓navigate Enteropen