Token Counter
The TokenCounter class has as its main goal to calculate the number of tokens present in a text, using encoders based on OpenAI's language models.
On this page
#Overview
The TokenCounter class has as its main goal to calculate the number of tokens present in a text, using encoders based on OpenAI's language models. This is especially useful for anyone working with APIs that have limits based on the number of tokens, such as GPT-3 and GPT-4, allowing better management of the cost and size of requests.
It solves the problem of needing an accurate token count for text inputs, something crucial for limits and optimizations in NLP applications that depend on OpenAI models. In practice, you can use this class to, for example, check whether a text is within the maximum token limit allowed by the model before sending a request.
#Execution Flow
- Instantiate the
TokenCounterclass, passing the name of the model you want to use for token encoding (e.g."gpt-3.5-turbo"). This sets the tokenization standard appropriate for the model. - If the model name is not recognized, the class automatically uses a default base encoding (
cl100k_base). - Call the
countmethod, passing the text you want to count. - The method returns an integer that represents the number of tokens in the text.
#Class Methods Table
| Method | Description |
|---|---|
__init__ | Initializes the encoder for the desired model |
count | Counts the number of tokens in a text |
#Environment Variables
No environment variables are used by the class.
#Key Architecture Points and Insights
- The class uses the
tiktokenlibrary, which is specific to tokenization of OpenAI models, ensuring high compatibility and accuracy. - The choice of encoding is dynamic, based on the given model name. If the model is not known, a safe default is used to avoid errors.
- It encapsulates the internal workings of the encoder, offering a simple interface to count tokens, making integration into larger projects easier.
- The class does not depend on other classes besides the external module
tiktoken. - Its design is minimalist, focused on a single purpose, following the single responsibility principle.
#Class and Methods Description
#TokenCounter Class
#Description
Class responsible for counting the number of tokens in texts, using encoders from the tiktoken library, which reproduce how OpenAI models interpret and split words into tokens.
#Constructor Arguments
| Argument | Type | Description | Default Value |
|---|---|---|---|
| model | str | Name of the OpenAI model to define the token encoding (e.g. "gpt-3.5-turbo"). | None |
#Methods
#1. __init__
Description
Initializes a token encoder based on the specified model. If the model is not recognized, it uses a default encoding.
Arguments
- model (str): name of the model to define the encoding
Returns
- Returns no value.
Raises
- None explicitly, but it ignores
KeyErrorwhen looking up the encoder for the model.
Examples
counter = TokenCounter("gpt-3.5-turbo")#2. count
Description
Counts how many tokens the given text has, converting the text using the encoder associated with the model.
Arguments
- text (str): Text whose number of tokens will be computed.
Returns
- int: total number of tokens found in the text. Returns 0 if the text is empty or
None.
Raises
- None.
Examples
counter = TokenCounter("gpt-3.5-turbo")
num_tokens = counter.count("Hello, how are you?")
print(num_tokens) # Saída provável: 6
# python -m src.tokens_calculate.token_counterThis example shows the approximate token count for the sentence "Hello, how are you?", which in general has 6 tokens according to the tokenization of the indicated model.
Source: src/tokens_calculate/token_counter.py