Vector Store

Pinecone Embedding

The PineconeEmbedding class is an implementation focused on making it easier to work with embedding vectors based on the Pinecone library, widely used to store and search vectors at scale.

On this page

    #Overview

    The PineconeEmbedding class is an implementation focused on making it easier to work with embedding vectors based on the Pinecone library, widely used to store and search vectors at scale. It offers features to split texts into smaller pieces, turn those pieces into structured documents with metadata, generate embedding vectors from the texts and manage the storage of those vectors in Pinecone indexes.

    This service is ideal for applications that deal with large volumes of text data and need to perform semantic search or embedding-based recommendations. Typical use involves processing raw text, generating embeddings through OpenAI, organized and efficient storage of those vectors, plus the ability to maintain them by removing specific vectors.

    #Execution Flow

    1. Class Initialization: The class is instantiated with a custom or default Pinecone client, configuring embedding models, chunk sizes, namespaces, and other configuration parameters loaded from the environment or files.
    2. Vector Generation:
      • Receives a full text and associated metadata (e.g. file_id).
      • The text is split into smaller pieces respecting size and overlap limits.
      • Each piece becomes a document with metadata.
      • The documents are turned into embedding vectors in batches.
      • The vectors are saved in Pinecone's main namespace, with optional saving to the global namespace for reuse.
    3. Vector Deletion:
      • Makes it possible to delete vectors from Pinecone by filtering on a metadata field and value, in the specified namespace.
      • Handles deletions in batches for efficiency.
      • Ensures rollback and logging to avoid inconsistencies in case of failures.
    4. Monitoring and Logs:
      • Integration with the tracing system (ApplicationTracing) for detailed debugging of the steps.
      • Careful exception handling to report errors that occurred in the process.

    #Class Methods Table

    MethodDescription
    __init__Initializes the service, configures client and parameters.
    embedding_documentCreates and saves embeddings of text split into chunks.
    delete_documentsDeletes vectors in Pinecone based on metadata filters.
    split_textSplits a long text into multiple pieces.
    build_documentsCreates Document objects with content and metadata. (static method)
    generate_embeddingsGenerates the vector representations from texts.

    #Environment Variables

    • OPENAI_EMBEDDING_MODEL: Defines the OpenAI embedding model to use to generate the vectors when none is explicitly provided.

    #Key Architecture Points and Insights

    • Inheritance and Reuse: PineconeEmbedding inherits from EmbeddingHelpers, separating the text handling and embedding generation logic from the higher-level storage and management flow.
    • Smart Batching: Embedding generation and deletion are processed in configurable batches to balance performance and resource use.
    • Separate Namespaces: The distinction between main and global namespace allows data organization, enabling isolated and shared scope scenarios.
    • Deep Logging and Tracing: Integration with a custom tracing system that enables granular debugging and detailed monitoring of operations.
    • Robust Error Handling: Explicit rollback on batch insertion failures keeps the Pinecone index consistent.
    • Use of OpenAIEmbeddings: Abstraction of embedding generation through an OpenAI model that can be easily swapped via environment variable or argument.
    • External Dependencies: The class depends on the Pinecone client implementation (PineconeClient), the configuration (PineconeVectorStoreConfig) and the LangChain library for handling documents and embeddings.

    #Class and Methods Description

    #Class PineconeEmbedding

    Description

    Class to manage the creation, storage and removal of text embeddings using Pinecone as the vector search service. It splits texts into pieces, generates embeddings through OpenAI, and lets you persist those vectors in configurable namespaces, easing integrations in systems that require semantic search.

    Constructor Arguments

    ArgumentTypeDescriptionDefault Value
    vector_clientOptional[PineconeClient]Custom Pinecone client for vector operations. If not provided, creates a default one.None
    embedding_model_namestrName of the embedding model used. If None, looks up the environment variable OPENAI_EMBEDDING_MODEL or the config default.None
    dimensionsintDimensionality of the vectors; if not given, uses the default configuration.None

    #1. __init__

    Description

    Initializes the embedding service, configuring the Pinecone client, parameters such as chunk size, namespaces and embedding model, and creates the global and main Pinecone stores.

    Arguments

    • vector_client (Optional[PineconeClient]): Pinecone client for operations.
    • embedding_model_name (str): embedding model name.
    • dimensions (int): vector dimension.

    Returns

    • None.

    Raises

    • RuntimeError: On failures during initialization.

    Examples

    service = PineconeEmbedding()
    # ou com cliente customizado e modelo definido
    service_custom = PineconeEmbedding(
        vector_client=my_pinecone_client,
        embedding_model_name="text-embedding-ada-002",
        dimensions=1536
    )

    #2. embedding_document

    Description

    Processes long text, splits it into chunks, creates documents associating metadata, generates embeddings for those chunks in batches, and saves the vectors in Pinecone. It can save copies in the global namespace for reuse.

    Arguments

    • text (str): full text to generate vectors from.
    • metadata (dict): metadata to associate with the vector documents.
    • save_global (bool): flag to also save the vectors in the global namespace.
    • batch_size (int | None): batch size for embedding generation.

    Returns

    • dict: Result with status, messages, and detailed information about the saved embeddings, including ids and count.

    Raises

    • Not directly; errors are caught and reported in the return value with rollback.

    Examples

    response = service.embedding_document(
        text="Um texto muito longo que precisa ser dividido e vetorizado...",
        metadata={"file_id": "doc_001"},
        save_global=True,
        batch_size=50
    )
    if response["status"] == "success":
        print(f"Vetores salvos, batches: {response['embedding_informations']['batch_count']}")
    else:
        print("Erro:", response["message"])

    #3. delete_documents

    Description

    Removes vectors from the Pinecone index by filtering on the value of a metadata field, within the specified namespace. It runs deletions in batches for performance and safety.

    Arguments

    • target_feature (str): metadata field used for filtering (e.g. "file_id").
    • target_id (str): value that identifies the vectors to delete.
    • namespace (str): Pinecone namespace in which to run the removal.
    • features (list): optional list of valid features for validating the filter.

    Returns

    • dict: Information about the number of deleted vectors and the namespace.

    Raises

    • ValueError: if target_feature is not in the features list when provided.
    • RuntimeError: if an error occurs during removal.

    Examples

    result = service.delete_documents(
        target_feature="file_id",
        target_id="doc_001",
        namespace="embedding_file"
    )
    print(f"Vectors deleted: {result['deleted_vectors']}")

    #4. split_text

    Description

    Splits a text into multiple smaller pieces, respecting size and overlap limits to ease processing and embedding generation.

    Arguments

    • text (str): text to be split.
    • chunk_size (int | None): desired chunk size (optional).
    • chunk_overlap (int | None): overlap between chunks (optional).

    Returns

    • List[str]: list with the generated text pieces.

    Raises

    • RuntimeError: if splitting the text fails.

    Examples

    chunks = service.split_text("Texto muito longo...", chunk_size=200, chunk_overlap=40)
    print(f"Quantidade de chunks gerados: {len(chunks)}")

    #5. build_documents (static method)

    Description

    Creates Document objects (from LangChain) out of text pieces, associating a metadata dictionary with each document.

    Arguments

    • chunks (List[str]): list of fragmented texts.
    • metadata (Dict[str, Any]): metadata to attach to each document.

    Returns

    • List[Document]: list of documents ready for ingestion.

    Raises

    • RuntimeError: if building the Documents fails.

    Examples

    documents = PineconeEmbedding.build_documents(
        ["texto 1", "texto 2"],
        {"file_id": "abc123", "created_at": "2024-06-01"}
    )
    for doc in documents:
        print(doc.metadata)

    #6. generate_embeddings

    Description

    Generates the vector embeddings for a list of texts, using the configured model.

    Arguments

    • texts (List[str]): list of sentences or documents to turn into vectors.

    Returns

    • List[List[float]]: list of numeric vectors matching the texts.

    Raises

    • RuntimeError: if an error occurs while generating the embeddings.

    Examples

    embeddings = service.generate_embeddings([
        "primeiro texto",
        "segundo texto"
    ])
    print(f"Embedding do primeiro texto tem dimensão {len(embeddings[0])}")

    #Full usage example

    pine_client = PineconeClient(index_name="exemplo-index", main_namespace="embeddings")
    service = PineconeEmbedding(vector_client=pine_client, embedding_model_name="text-embedding-ada-002")
    
    texto = "Este é um exemplo de texto que será dividido, embeddado e armazenado."
    
    response = service.embedding_document(
        text=texto,
        metadata={"file_id": "exemplo_12345"},
        save_global=True
    )
    
    if response['status'] == "success":
        print("Embeddings gerados e salvos com sucesso!")
    else:
        print("Falha ao salvar embeddings:", response["message"])

    This documentation details how PineconeEmbedding works, highlighting its practical use, the available methods and tips for efficient integration with vectorization systems based on Pinecone and OpenAI.

    Source: src/vector_store/pinecone/embedding.py

    Esc
    ↑↓navigate Enteropen