Tokenizer Lab

Explore Connor Love’s tokenizer alongside nine established tokenizer baselines.

10tokenizers
compared

Token count by tokenizer

0 / 10 measured

Results

Connor’s Tokenizer

Byte-lossless Unigram · research candidate

loading
Method research candidateTokenizer byte-lossless Unigram · 195,124 vocabularyPin production-corpus candidateSource

OpenAI

o200k_base

loading
Method official tokenizerTokenizer tiktoken byte-level BPE · 200K vocabularyPin js-tiktoken@1.0.21Source

Anthropic

Anthropic legacy tokenizer · local proxy

loading
Method official tokenizerTokenizer Anthropic 2023 BPE · pre-Claude 3Pin @anthropic-ai/tokenizer@0.0.4Source

Google

Gemma 4 12B · open model

loading
Method official tokenizerTokenizer Gemma BPE · 262,144 vocabulary · not GeminiPin 023679eSource

Moonshot AI

Kimi K2.6 · open fallback

loading
Method official tokenizerTokenizer TikToken BPE · 163,584 vocabulary · K3 pendingPin 81bcaaaSource

DeepSeek

DeepSeek-V3.2

loading
Method official tokenizerTokenizer LlamaTokenizerFast BPE · 128K vocabularyPin c69397eSource

xAI

Grok-1 · open local proxy

loading
Method official tokenizerTokenizer SentencePiece · 131,072 vocabulary · not Grok 4.5Pin Xenova conversion@40ee9aeSource

Meta

Llama 3.1 tokenizer family

loading
Method official tokenizerTokenizer TikToken-based BPE · 128K vocabulary · converted assetPin 72bff9eSource

Mistral AI

Mistral Small 4

loading
Method official tokenizerTokenizer Tekken BPE · 131,072 vocabularyPin 233ef01Source

Alibaba Cloud

Qwen3

loading
Method official tokenizerTokenizer byte-level BPE · 151,643 vocabularyPin b968826Source