tiktoken: Count LLM Tokens in Python (2026)
tiktoken: Count LLM Tokens in Python (2026) shows how to measure prompts the way OpenAI models actually see them: pick the right encoding, count tokens, inspect how text is split, truncate to a budget, chunk documents with overlap for RAG, and stop special-token surprises before they reach the API. Every number below comes from a real run.
Chunked text feeds straight into a vector store like ChromaDB or LanceDB hybrid search, and token budgets matter even more when you stream LLM tokens with FastAPI SSE.
TL;DR
tiktoken.encoding_for_model("gpt-4o")returns theo200k_baseencoding used by current OpenAI models.len(enc.encode(text))is your token count;enc.decode(ids)round-trips back to the exact text.- Truncate and chunk on token boundaries, not characters, so you never overflow a context window.
- Strings like
<|endoftext|>raiseValueErrorby default; passdisallowed_special=()to treat user text as plain text. enc.encode_batch()counts many prompts in one call. We tested tiktoken0.14.0on Python3.13.5.
Why count tokens in 2026?
LLM APIs bill per token, cap requests by a context window measured in tokens, and rate-limit by tokens per minute. Character counts are a poor guess: English averages roughly four characters per token, but code, numbers, and non-Latin scripts behave very differently. Counting locally with tiktoken is fast, free, and works offline after the first download of the encoding file.
| API | What it does | Typical use |
|---|---|---|
encoding_for_model() | Model name to encoding | Match the tokenizer to your model |
get_encoding("o200k_base") | Load an encoding by name | Pin the tokenizer explicitly |
encode() / decode() | Text to token ids and back | Count, truncate, chunk |
encode_batch() | Encode many strings at once | Pre-flight budget checks |
Versions tested (2026-10-05)
- Python
3.13.5 tiktoken0.14.0- Encodings:
o200k_base(gpt-4o and newer) andcl100k_base(older gpt-4 / gpt-3.5 era)
python -m venv .venv && source .venv/bin/activate
pip install tiktoken==0.14.0
python token_demo.py
The first call downloads the encoding file and caches it. Set TIKTOKEN_CACHE_DIR to a folder you ship with your app if production servers have no internet access.
1. Count tokens for a model
import sys
import tiktoken
print("python", sys.version.split()[0], "| tiktoken", tiktoken.__version__)
# 1) Pick the encoding for a model, then count tokens
enc = tiktoken.encoding_for_model("gpt-4o")
text = "Python makes LLM apps easy. Count tokens before you call the API!"
ids = enc.encode(text)
print("encoding:", enc.name)
print("chars:", len(text), "| tokens:", len(ids))
python 3.13.5 | tiktoken 0.14.0
encoding: o200k_base
chars: 65 | tokens: 15
65 characters became 15 tokens. Always choose the encoding from the model name so your counts match what the API bills.
2. See how text is split
# 2) See how the text is split
pieces = [enc.decode([t]) for t in ids[:8]]
print("first pieces:", pieces)
print("round trip ok:", enc.decode(ids) == text)
first pieces: ['Python', ' makes', ' L', 'LM', ' apps', ' easy', '.', ' Count']
round trip ok: True
Notice that LLM splits into ' L' and 'LM', and leading spaces belong to the next token. That is why slicing text by characters can cut a word in half while slicing token ids never breaks the round trip.
3. Compare encodings (English vs Urdu)
# 3) Same text, two encodings
old = tiktoken.get_encoding("cl100k_base")
urdu = "پائتھون سیکھنا آسان ہے"
for label, s in [("english", text), ("urdu", urdu)]:
print(f"{label}: cl100k={len(old.encode(s))} o200k={len(enc.encode(s))}")
english: cl100k=15 o200k=15
urdu: cl100k=24 o200k=10
English costs the same in both encodings, but the Urdu sentence drops from 24 tokens to 10 with o200k_base. If your users write in Urdu, Hindi, Arabic, or other non-Latin scripts, the newer encoding means cheaper calls and more room in the context window.
4. Truncate a prompt to a token budget
# 4) Truncate a long prompt to a token budget
def truncate(s: str, max_tokens: int) -> str:
toks = enc.encode(s)
return s if len(toks) <= max_tokens else enc.decode(toks[:max_tokens])
doc = " ".join(f"Line {i}: tiktoken counts tokens fast." for i in range(1, 201))
short = truncate(doc, 50)
print("doc tokens:", len(enc.encode(doc)), "| truncated:", len(enc.encode(short)))
doc tokens: 2200 | truncated: 50
5. Chunk documents with overlap for RAG
# 5) Chunk by tokens with overlap (for RAG / embeddings)
def chunk(s: str, size: int = 200, overlap: int = 20) -> list[str]:
toks = enc.encode(s)
step = size - overlap
return [enc.decode(toks[i:i + size]) for i in range(0, len(toks), step)]
chunks = chunk(doc)
print("chunks:", len(chunks), "| sizes:", [len(enc.encode(c)) for c in chunks])
chunks: 13 | sizes: [200, 200, 200, 200, 200, 200, 200, 200, 200, 200, 200, 200, 40]
Each chunk is exactly 200 tokens with a 20-token overlap so a sentence cut at a boundary still appears whole in the next chunk. Feed these chunks to your embedding model and vector store.
6. Handle special tokens safely
# 6) Special tokens are blocked by default
try:
enc.encode("user text <|endoftext|>")
except ValueError:
print("special token blocked: ValueError")
print("allowed as text:", len(enc.encode("user text <|endoftext|>", disallowed_special=())))
special token blocked: ValueError
allowed as text: 9
By default tiktoken refuses to encode control strings like <|endoftext|>. That protects you from prompt-injection tricks, but it also crashes when a user pastes that text. Passing disallowed_special=() encodes it as ordinary characters.
7. Budget check a batch before sending
# 7) Budget check before you send a batch
budget = 1000
prompts = [text, doc, short]
counts = enc.encode_batch(prompts)
for p, c in zip(["question", "doc", "short"], counts):
print(f"{p:9} {len(c):5} tokens fits={len(c) <= budget}")
question 15 tokens fits=True
doc 2200 tokens fits=False
short 50 tokens fits=True
Full real output
python 3.13.5 | tiktoken 0.14.0
encoding: o200k_base
chars: 65 | tokens: 15
first pieces: ['Python', ' makes', ' L', 'LM', ' apps', ' easy', '.', ' Count']
round trip ok: True
english: cl100k=15 o200k=15
urdu: cl100k=24 o200k=10
doc tokens: 2200 | truncated: 50
chunks: 13 | sizes: [200, 200, 200, 200, 200, 200, 200, 200, 200, 200, 200, 200, 40]
special token blocked: ValueError
allowed as text: 9
question 15 tokens fits=True
doc 2200 tokens fits=False
short 50 tokens fits=True
Common mistakes
- Using the wrong encoding.
cl100k_basecounts differ fromo200k_base, especially for non-English text. - Counting characters. Use
len(enc.encode(text)), neverlen(text) / 4, when the limit matters. - Forgetting the reply. The context window covers input and output, so leave room for the model's answer.
- Assuming one tokenizer fits all. Claude, Gemini, and Llama use their own tokenizers; treat tiktoken counts as an estimate for them.
FAQ
Does tiktoken call the OpenAI API? No. It runs locally; it only downloads the encoding file once and caches it.
Which encoding should I use for GPT-4o and newer models? o200k_base, which encoding_for_model("gpt-4o") returns in our test.
How do I estimate cost? Multiply the token count by your provider's current per-million-token price from its pricing page, and add the expected output tokens.
Next steps
Wrap truncate() and chunk() in a small utility module, call it before every LLM request, and log the token count next to each response so you can track cost and latency over time.