Vector Store

Pinecone Client

The PineconeClient class is an abstraction for managing the connection and operations with the Pinecone service, a high-performance vector storage and search solution.

On this page

    #Overview

    The PineconeClient class is an abstraction for managing the connection and operations with the Pinecone service, a high-performance vector storage and search solution. It makes integrating with external APIs (Pinecone and OpenAI) easier, taking care of API key configuration, namespace definition and initialization of the embedding model to work with text vectors.

    This class solves the problem of manual, repetitive configuration of these integrations, providing a unified and safe interface to create and manipulate vector stores using custom namespaces and specific embedding models. In practice, it can be used to store and search vectors generated from texts, which is fundamental for systems such as semantic search, similarity-based recommendation and NLP.

    #Execution Flow

    1. Client Initialization:
      • When an instance of PineconeClient is created, the OpenAI and Pinecone API keys are loaded from the environment variables.
      • index_name, main_namespace and global_namespace are defined either through the constructor parameters or through environment variables or the default configuration.
      • The embedding model to use is selected the same way.
      • It internally initializes the connection with Pinecone and the OpenAI embedding model.
    2. Connection Setup:
      • _init_pinecone creates the Pinecone client and instantiates the specified index.
      • _init_embeddings initializes the OpenAI embedding model according to the configuration.
    3. Namespace Resolution:
      • The get_namespace() method returns the namespace to use in operations, giving priority to a value passed to the method, if provided, or else the instance's main namespace.
    4. Vector Store Creation:
      • create_vector_store() builds the PineconeVectorStore instance, associating the appropriate index and embedding model with the desired namespace.
      • The returned object allows vector similarity operations.

    #Class Methods Table

    MethodDescription
    __init__Initializes the Pinecone client, loading settings and keys
    _init_pineconeEstablishes the connection with Pinecone and instantiates the index
    _init_embeddingsInitializes the OpenAI embedding model
    get_namespaceReturns the namespace resolved for operations
    create_vector_storeReturns a configured instance of PineconeVectorStore

    #Environment Variables

    • OPENAI_API_KEY: OpenAI API key for generating embeddings.
    • PINECONE_API_KEY: Pinecone API key for accessing the service.
    • PINECONE_INDEX_NAME: Name of the Pinecone index to use.
    • PINECONE_NAMESPACE: Default main namespace for storing vectors.
    • PINECONE_GLOBAL_NAMESPACE: Optional global namespace for shared vectors.
    • OPENAI_EMBEDDING_MODEL: Name of the OpenAI embedding model to use.

    #Key Architecture Points and Insights

    • The class uses encapsulation to hide the initialization details of Pinecone and of the embedding model.
    • Using environment variables with a fallback to constructor parameters increases flexibility and reuse.
    • The "dependency injection" pattern is present in the creation of the vector store, allowing embedding and namespace to be overridden in the call.
    • Integration with specialized external modules (langchain_openai, langchain_pinecone), showing the use of composition.
    • The class uses a custom tracing system (ApplicationTracing) to ease debugging and monitoring, producing detailed logs at every step.
    • Clear separation between configuration, initialization and vector store creation, favoring maintenance and testing.

    #Class and Methods Description

    #Class PineconeClient

    Description

    Class for controlling and easing the connection with the Pinecone service and for managing the stored vectors. It integrates the Pinecone and OpenAI APIs to create a custom vector base, which can be queried or updated from the embedding models provided. It allows namespaces and models to be configured flexibly, simplifying vector operations in applications.

    Constructor Arguments

    ArgumentTypeDescriptionDefault Value
    index_nameOptional[str]Name of the Pinecone index to useNone
    main_namespaceOptional[str]Primary namespace for storing vectorsNone
    global_namespaceOptional[str]Optional global namespace for shared vectorsNone
    embedding_modelOptional[str]Name of the OpenAI model for embeddingsNone

    #1. __init__

    Description

    Initializes the Pinecone client by loading the API keys and settings, defining the namespaces and the embedding model, and establishing connections.

    Arguments

    • index_name (Optional[str]): Pinecone index name (optional)
    • main_namespace (Optional[str]): primary namespace (optional)
    • global_namespace (Optional[str]): global namespace (optional)
    • embedding_model (Optional[str]): OpenAI model name (optional)

    Returns

    • Does not return a value.

    Raises

    • EnvironmentError: if the API keys are not defined.
    • ValueError: if index_name is not defined.
    • RuntimeError: for general initialization failures.

    Examples

    client = PineconeClient(index_name="meu_indice", main_namespace="app_namespace")

    #2. _init_pinecone

    Description

    Establishes the connection with the Pinecone service and instantiates the index with the configured name.

    Arguments

    • None.

    Returns

    • Does not return a value.

    Examples

    client._init_pinecone()
    # Internamente conecta ao Pinecone e configura o index para operações

    #3. _init_embeddings

    Description

    Initializes the OpenAI model for generating embeddings, optionally using a custom model name.

    Arguments

    • model_name (Optional[str]): embedding model name (optional)

    Returns

    • Does not return a value.

    Examples

    client._init_embeddings("text-embedding-ada-002")
    # Atualiza o modelo de embedding usado pelo cliente

    #4. get_namespace

    Description

    Returns the namespace to use for operations, giving priority to the argument passed, or else returning the configured main one.

    Arguments

    • namespace (Optional[str]): optional replacement namespace.

    Returns

    • str: namespace resolved for use.

    Examples

    ns = client.get_namespace()  # retorna o main_namespace configurado
    ns2 = client.get_namespace("namespace_alternativo")  # retorna "namespace_alternativo"

    #5. create_vector_store

    Description

    Creates and returns a PineconeVectorStore instance configured with the selected index, embedding and namespace.

    Arguments

    • namespace (Optional[str]): namespace for the store (optional).
    • embedding_model (Optional[OpenAIEmbeddings]): embedding model instance (optional).

    Returns

    • PineconeVectorStore: object for manipulating the vector storage.

    Examples

    vector_store = client.create_vector_store()
    # Usa main_namespace e modelo configurado
    
    vector_store_custom = client.create_vector_store(namespace="ns_custom")
    # Cria store com namespace customizado

    #Usage

    if __name__ == "__main__":
        client = PineconeClient()
        vector_store = client.create_vector_store()
        print("Pinecone Client initialized and VectorStore created successfully.")
    
    # python -m src.vector_store.pinecone.client

    Source: src/vector_store/pinecone/client.py

    Esc
    ↑↓navigate Enteropen