Lipește text pentru estimare Explain how HTTP caching works to a junior developer. Cover ETags, Cache-Control, and the Vary header in under 200 words.
(Documentație în engleză)
What it does
The Token Estimator gives you a fast, approximate token count for any text, code, or JSON before you send it to an LLM. Instead of running a real tokenizer, it classifies each line as prose, code, JSON, or CJK text and divides by that content type’s typical characters-per-token rate (4 / 3.5 / 3 / 1.5), then reports a point estimate with a ±15% band — because real BPE tokenizers vary by vocabulary and language mix. It can also add chat-framing overhead (~5 tokens per message) for role markers and delimiters. Everything runs 100% client-side: nothing you paste is ever sent to a server, so it is safe for sensitive prompts.
Need an exact count? Toggle Exact count and the real BPE tokenizer (OpenAI’s o200k_base encoding, the same vocabulary GPT-4o and later models use) loads on demand and counts your text precisely — exact for OpenAI-family models, a very close approximation for others (Anthropic and Google use different but similar-sized vocabularies). The tokenizer table (~2 MB) downloads only when you toggle it, and it too runs entirely in your browser. From either mode, Price a month of this → hands your token count straight to the LLM Cost Calculator’s monthly view.
How to use it
- Paste your text or code into the Text or code box (or hit Sample to load a demo prompt).
- Leave Content type on Auto to detect each line, or force Prose / Code / JSON / CJK to apply one rate to the whole input.
- Set Chat messages to the number of messages in your conversation — each adds 5 framing tokens to the total.
- Read the Estimated tokens headline, the low–high range, chars / words / lines, and the resolved content type.
- Copy tokens to grab the number, or Copy share link to send the exact input and settings to a teammate.
Examples
Plain prose
Input: Hello world → 3 tokens (≈ 3–3, type: prose)
Code
Input: const x = 1; → 3 tokens (type: code)
JSON gets the dense rate automatically
Input: {"a": 1, "b": 2} → 5 tokens (type: json — the whole document parses as JSON, so every line uses the 3 chars/token rate)
A longer block with a visible band
Input: 100 × a → 25 tokens (≈ 21–29)
Good to know
- These are estimates, not exact counts in the default mode: every result carries a ±15% band. Treat the number as a budgeting aid, not an invoice — model providers count tokens with model-specific vocabularies. Toggle Exact count for the real o200k_base tokenizer when the number has to be right.
- Auto mode detects JSON documents as a whole: if your input parses as valid JSON, every line uses the JSON rate. Force a content type to override that.
- CJK text is expensive: Chinese, Japanese, and Korean pack roughly one token per 1.5 characters — the same “page” of text can cost 2-3× more tokens than English.
- Shareable: your input and options are encoded in the URL (
?t=...&type=...&msg=...), so a link reproduces the exact estimate. - Private: all estimation happens locally - safe for secrets and credentials.
- Related tools: Word Counter, Secure Token Generator, JSON Formatter.