Token Calculate

Token Counter

The TokenCounter class has as its main goal to calculate the number of tokens present in a text, using encoders based on OpenAI's language models.

On this page

    #Overview

    The TokenCounter class has as its main goal to calculate the number of tokens present in a text, using encoders based on OpenAI's language models. This is especially useful for anyone working with APIs that have limits based on the number of tokens, such as GPT-3 and GPT-4, allowing better management of the cost and size of requests.

    It solves the problem of needing an accurate token count for text inputs, something crucial for limits and optimizations in NLP applications that depend on OpenAI models. In practice, you can use this class to, for example, check whether a text is within the maximum token limit allowed by the model before sending a request.

    #Execution Flow

    1. Instantiate the TokenCounter class, passing the name of the model you want to use for token encoding (e.g. "gpt-3.5-turbo"). This sets the tokenization standard appropriate for the model.
    2. If the model name is not recognized, the class automatically uses a default base encoding (cl100k_base).
    3. Call the count method, passing the text you want to count.
    4. The method returns an integer that represents the number of tokens in the text.

    #Class Methods Table

    MethodDescription
    __init__Initializes the encoder for the desired model
    countCounts the number of tokens in a text

    #Environment Variables

    No environment variables are used by the class.

    #Key Architecture Points and Insights

    • The class uses the tiktoken library, which is specific to tokenization of OpenAI models, ensuring high compatibility and accuracy.
    • The choice of encoding is dynamic, based on the given model name. If the model is not known, a safe default is used to avoid errors.
    • It encapsulates the internal workings of the encoder, offering a simple interface to count tokens, making integration into larger projects easier.
    • The class does not depend on other classes besides the external module tiktoken.
    • Its design is minimalist, focused on a single purpose, following the single responsibility principle.

    #Class and Methods Description

    #TokenCounter Class

    #Description

    Class responsible for counting the number of tokens in texts, using encoders from the tiktoken library, which reproduce how OpenAI models interpret and split words into tokens.

    #Constructor Arguments

    ArgumentTypeDescriptionDefault Value
    modelstrName of the OpenAI model to define the token encoding (e.g. "gpt-3.5-turbo").None

    #Methods

    #1. __init__

    Description

    Initializes a token encoder based on the specified model. If the model is not recognized, it uses a default encoding.

    Arguments

    • model (str): name of the model to define the encoding

    Returns

    • Returns no value.

    Raises

    • None explicitly, but it ignores KeyError when looking up the encoder for the model.

    Examples

    counter = TokenCounter("gpt-3.5-turbo")

    #2. count

    Description

    Counts how many tokens the given text has, converting the text using the encoder associated with the model.

    Arguments

    • text (str): Text whose number of tokens will be computed.

    Returns

    • int: total number of tokens found in the text. Returns 0 if the text is empty or None.

    Raises

    • None.

    Examples

    counter = TokenCounter("gpt-3.5-turbo")
    num_tokens = counter.count("Hello, how are you?")
    print(num_tokens)  # Saída provável: 6
    
    # python -m src.tokens_calculate.token_counter

    This example shows the approximate token count for the sentence "Hello, how are you?", which in general has 6 tokens according to the tokenization of the indicated model.

    Source: src/tokens_calculate/token_counter.py

    Esc
    ↑↓navigate Enteropen