Vector Store / Local Dynamic Embedding

Documentation

The LocalDynamicEmbedding class offers a local pipeline for text processing, where the text is split into pieces (chunks), each piece is converted into an embedding vector, and those vectors are stored for later retrieval by similarity.

On this page

    #Overview

    The LocalDynamicEmbedding class offers a local pipeline for text processing, where the text is split into pieces (chunks), each piece is converted into an embedding vector, and those vectors are stored for later retrieval by similarity. This pipeline supports different embedding providers, including real providers (such as OpenAI and Huggingface) and fake embeddings for testing, making experimentation easier without the need for external APIs.

    The main problem solved is working with long texts for retrieval and search tasks without depending exclusively on external services, optimizing text splitting and embedding computation in a modular and configurable way. In addition, it provides a fluent API that lets you configure step by step the embedding provider, the text splitter, and the number of results retrieved.

    In practice, it can be used to build local document search systems, chatbots that need to understand documents or knowledge bases, and for quick prototyping, where you want to easily swap the embedding provider or the chunking parameters.

    #Execution Flow

    1. Pipeline configuration: Initially, the user creates an instance of LocalDynamicEmbedding and configures the necessary components using the fluent model:
      • Sets the embedding provider with with_provider() or convenient methods such as from_openai_embeddings().
      • Adjusts the text splitter with with_splitter().
      • Configures how many results to retrieve with with_top_k().
    2. Text processing: The process_text(texto) method receives an input text, splits it into chunks based on the defined parameters, generates embeddings for each chunk and stores everything in an in-memory FAISS vector index.
    3. Information retrieval: With the index created, retrieve(query) is called to look up the chunks most similar to the query, returning the texts, metadata, similarity scores, and optionally the embedding vectors.
    4. Managing and querying the chunks: The user can access the stored chunks via properties and methods such as chunks, get_chunks() or get_chunk(index) to explore the content and the vectors.
    5. Reconfiguration and cleanup: If you want to change the configuration or reset the pipeline, clear() clears the state, allowing reconfiguration without creating a new instance.

    #Class Methods Table

    MethodDescription
    __init__Constructor that initializes parameters and state.
    with_embeddingsSets a custom embeddings instance.
    with_fake_embeddingsConfigures the use of fake embeddings.
    with_providerConfigures the embeddings provider by name.
    with_splitterAdjusts the splitter parameters (chunk size, overlap).
    with_top_kSets the number of results in retrieval.
    from_providerInstantiates the class with an embeddings provider.
    from_openai_embeddingsCreates an instance with OpenAI embeddings.
    from_huggingface_embeddingsCreates an instance with Huggingface embeddings.
    from_fake_embeddingsCreates an instance with fake embeddings.
    process_textProcesses the text, splitting, computing embeddings and storing.
    retrieveRetrieves the most relevant chunks for a query.
    as_retrieverReturns an object to run external searches.
    chunksProperty that returns the list of stored chunks.
    get_chunksReturns the chunks as dictionaries.
    get_chunkReturns a specific chunk by index.
    total_chunksReturns the total number of stored chunks.
    clearClears the state for a new configuration or reuse.

    #Key Architecture Points and Insights

    • Builder Pattern / Fluent API: The class adopts a builder pattern for fluent configuration, allowing the user to configure the pipeline step by step before processing, which increases flexibility in building the flow.
    • Lazy Initialization: Components such as embeddings and splitter are initialized only when needed, allowing changes and configuration before processing takes effect.
    • Separation of Responsibilities: Embedding creation is delegated to EmbeddingFactory, which abstracts different providers (OpenAI, Huggingface, fake), making extension and maintenance easier.
    • Local Indexing with FAISS: Stores the embeddings in a local FAISS index for fast in-memory search, which is efficient for prototypes and smaller-scale applications without external infrastructure.
    • State Management and Guard Rails: After processing (process_text), reconfiguration of the main components is blocked to avoid inconsistencies, and clear() allows resetting the pipeline.
    • Helper Class Chunk: Represents each piece of the text with its content, index, metadata and embedding vector, making it easier to access the information in a structured way.
    • Complete Chunk Inspection: Each chunk gives easy access to its text, metadata, embedding vector, length in characters and embedding dimension, meeting debugging and analysis needs.
    • Use of .env and dotenv: Loads environment variables for authentication (e.g. OpenAI), although that responsibility stays with the embedding libraries.

    #Class and Methods Description

    #Class LocalDynamicEmbedding

    Description

    Class to assemble a local pipeline for chunking, embedding generation and retrieval via similarity search. It supports multiple embedding providers configured through a fluent API, lets you adjust how the text is split and how the results are selected, operating locally with FAISS for indexing and search.

    It allows processing long texts by splitting them into pieces, computing embeddings for each piece, storing those vectors and retrieving them with queries, all in a modular and reusable way.

    Constructor Arguments

    ArgumentTypeDescriptionDefault Value
    embeddingsOptional[Embeddings]Custom instance of the embeddings object to use.None (uses internal fake)
    sizeintEmbedding vector dimension (used for fake embeddings).384
    chunk_sizeintMaximum number of characters per piece when splitting the text.500
    chunk_overlapintAmount of overlap between pieces to keep context.50
    top_kintNumber of top results to return in retrieval.4
    separatorsOptional[List[str]]List of separators used to split the text (e.g. "\n\n", "\n", etc.)Defined default list

    Methods

    #1. __init__

    Description

    Initializes the pipeline with the chunking, embeddings and retrieval parameters, setting the initial state and preparing attributes for lazy loading.

    Arguments

    • embeddings (Optional[Embeddings]): Custom embeddings instance, optional.
    • size (int): Embedding vector size for fake embeddings.
    • chunk_size (int): Maximum chunk size.
    • chunk_overlap (int): Overlap between chunks.
    • top_k (int): Number of results to return.
    • separators (Optional[List[str]]): Separators used to split the text.

    Returns

    • Does not return a value (None).

    Raises

    • Not specified.

    Examples

    pipeline = LocalDynamicEmbedding(chunk_size=1000, top_k=3)

    #2. with_embeddings

    Description

    Sets a custom embeddings instance for the pipeline before any processing.

    Arguments

    • embeddings (Embeddings): Embeddings instance to use.

    Returns

    • LocalDynamicEmbedding: Returns self to allow call chaining.

    Raises

    • RuntimeError: If called after processing has started or with invalid embeddings.

    Examples

    pipeline = LocalDynamicEmbedding().with_embeddings(custom_embeddings)

    #3. with_fake_embeddings

    Description

    Configures the pipeline to use fake embeddings, with the option of setting the vector size.

    Arguments

    • size (Optional[int]): Dimension size of the fake embeddings.

    Returns

    • LocalDynamicEmbedding: Returns self for chaining.

    Raises

    • RuntimeError: If the size is invalid or an internal error occurs.

    Examples

    pipeline = LocalDynamicEmbedding().with_fake_embeddings(size=128)

    #4. with_provider

    Description

    Configures the embeddings provider by name and specific parameters (e.g. "openai", "huggingface", "fake").

    Arguments

    • provider (str): Provider name.
    • **kwargs: Additional parameters for the provider.

    Returns

    • LocalDynamicEmbedding: Returns self for chaining.

    Raises

    • RuntimeError: If the provider cannot be configured or instantiated.

    Examples

    pipeline = LocalDynamicEmbedding().with_provider("openai", model="text-embedding-3-large")

    #5. with_splitter

    Description

    Configures the text splitter parameters: chunk size, overlap between chunks and separators.

    Arguments

    • chunk_size (Optional[int]): Maximum chunk size.
    • chunk_overlap (Optional[int]): Overlap between chunks.
    • separators (Optional[List[str]]): Separators used for splitting.

    Returns

    • LocalDynamicEmbedding: Returns self for chaining.

    Raises

    • RuntimeError: If called after text processing has started.

    Examples

    pipeline = LocalDynamicEmbedding().with_splitter(chunk_size=1000, chunk_overlap=100)

    #6. with_top_k

    Description

    Sets how many top results will be returned in searches.

    Arguments

    • top_k (int): Number of top results.

    Returns

    • LocalDynamicEmbedding: Returns self for chaining.

    Raises

    • No.

    Examples

    pipeline = LocalDynamicEmbedding().with_top_k(10)

    #7. from_provider

    Description

    Class member that creates a preconfigured instance for any provider with splitter and top_k parameters.

    Arguments

    • provider (str): Provider name.
    • chunk_size (int): Chunk size.
    • chunk_overlap (int): Chunk overlap.
    • top_k (int): Number of results to return.
    • separators (Optional[List[str]]): Separators for splitting.
    • **provider_kwargs: Arguments for creating the provider.

    Returns

    • LocalDynamicEmbedding: Configured instance.

    Raises

    • Not documented.

    Examples

    pipeline = LocalDynamicEmbedding.from_provider("huggingface", model="all-MiniLM-L6-v2")

    #8. from_openai_embeddings

    Description

    Creates an instance with OpenAI embeddings already configured; the model and chunk parameters can be customized.

    Arguments

    • model (str): OpenAI model.
    • chunk_size (int): Chunk size.
    • chunk_overlap (int): Chunk overlap.
    • top_k (int): Number of results.
    • separators (Optional[List[str]]): Separators.
    • **kwargs: Extra parameters for the provider.

    Returns

    • LocalDynamicEmbedding: OpenAI-configured instance.

    Examples

    pipeline = LocalDynamicEmbedding.from_openai_embeddings(model="text-embedding-3-large")

    #9. from_huggingface_embeddings

    Description

    Creates an instance configured to use Huggingface embeddings.

    Arguments

    • model (str): Huggingface model.
    • chunk_size (int), chunk_overlap (int), top_k (int), separators (Optional[List[str]]), **kwargs.

    Returns

    • LocalDynamicEmbedding: Huggingface-configured instance.

    Examples

    pipeline = LocalDynamicEmbedding.from_huggingface_embeddings()

    #10. from_fake_embeddings

    Description

    Creates an instance for using fake embeddings, with a parameterized size.

    Arguments

    • size (int): Size of the fake embeddings.
    • chunk_size (int), chunk_overlap (int), top_k (int), separators (Optional[List[str]]).

    Returns

    • LocalDynamicEmbedding: Instance configured with fake embeddings.

    Examples

    pipeline = LocalDynamicEmbedding.from_fake_embeddings(size=256)

    #11. process_text

    Description

    Processes a text, splitting it into chunks, computing embeddings, storing the chunks and adding them to the FAISS index.

    Arguments

    • text (str): Text to process (cannot be empty).
    • metadata (Optional[dict]): Metadata associated with the text.

    Returns

    • int: Number of chunks created.

    Raises

    • RuntimeError: If the text is empty or on a processing error.

    Examples

    num_chunks = pipeline.process_text("Texto extenso a ser processado", metadata={"source": "documento"})
    print(f"{num_chunks} chunks criados.")

    #12. retrieve

    Description

    Runs a similarity search on the index, returning the best chunks for a query.

    Arguments

    • query (str): Query for the search.
    • top_k (Optional[int]): Maximum number of results (overrides the one set).
    • include_embedding (bool): Whether to include embedding vectors in the result.

    Returns

    • List[Dict]: List of results with content, score, metadata, and optionally embedding.

    Raises

    • RuntimeError: If no text has been processed yet.

    Examples

    resultados = pipeline.retrieve("consulta de teste", top_k=3, include_embedding=True)
    for r in resultados:
        print(r["content"], r["score"])

    #13. as_retriever

    Description

    Returns a VectorStoreRetriever object configured to run queries on the index.

    Arguments

    • **kwargs: Additional parameters for configuring the retriever.

    Returns

    • VectorStoreRetriever: Object to run queries.

    Raises

    • RuntimeError: If the pipeline has not been processed.

    Examples

    retriever = pipeline.as_retriever()
    docs = retriever.get_relevant_documents("consulta")

    #14. chunks

    Description

    Property that returns the list with all the stored Chunk objects.

    Arguments

    • None.

    Returns

    • List[Chunk]: List of chunks.

    Examples

    for chunk in pipeline.chunks:
        print(chunk.content)

    #15. get_chunks

    Description

    Returns all chunks as dictionaries, with the option of including the embeddings.

    Arguments

    • include_embedding (bool): Defines whether the embedding vector should be included.

    Returns

    • List[dict]: List of dictionaries representing chunks.

    Examples

    chunks_data = pipeline.get_chunks(include_embedding=False)

    #16. get_chunk

    Description

    Gets a specific chunk by index.

    Arguments

    • index (int): Index of the desired chunk.

    Returns

    • Optional[Chunk]: The matching chunk, or None if it does not exist.

    Examples

    chunk = pipeline.get_chunk(0)
    print(chunk.content if chunk else "Chunk não encontrado")

    #17. total_chunks

    Description

    Property that returns the total of processed and stored chunks.

    Arguments

    • None.

    Returns

    • int: Total number of chunks.

    #18. clear

    Description

    Clears the whole state of the pipeline, including chunks and index, allowing reconfiguration.

    Arguments

    • None.

    Returns

    • LocalDynamicEmbedding: Returns self for chaining.

    Examples

    pipeline.clear()

    #Helper Class Chunk

    Description

    Represents a piece (chunk) of the original text with its content, metadata and embedding vector, allowing easy inspection of the data generated in the pipeline.

    Constructor Arguments

    ArgumentTypeDescriptionDefault Value
    indexintIndex of the chunk in the original text
    contentstrText content of the chunk
    metadatadictAssociated metadata
    embeddingList[float]Embedding vector of the chunk[]

    Properties

    • length: Returns the number of characters in the chunk's content.
    • dim: Returns the dimension of the embedding vector.

    Methods

    • to_dict(include_embedding=True): Returns a dictionary representing the chunk, optionally including embeddings.

    Examples

    chunk = Chunk(0, "Olá mundo", {"source": "doc1"}, [0.1, 0.2, 0.3])
    print(chunk.length)  # 9
    print(chunk.to_dict(include_embedding=False))  # {'index': 0, 'content': 'Olá mundo', ...}

    In this way, LocalDynamicEmbedding offers a modular, configurable and efficient pipeline for handling local embeddings on text, ideal for developers who want full control of the preprocessing and retrieval flow, as well as quick switching between embedding providers and adjustments to the chunking parameters.

    Source: src/embedding/modules/local_embedding/module.py

    Esc
    ↑↓navigate Enteropen