• Skip to primary navigation
  • Skip to main content
  • Skip to primary sidebar
  • Skip to footer
Sq Magazine LogoSQ Magazine

Smarter Insights for a Fast-Moving Digital World

  • Latest News
  • Statistics
  • About
  • Contact
  • Subscribe
Subscribe
Home » Glossary » A

What Is a Token in AI? How Models Count Text and Cost

Published on: August 13, 2026
Barry Elad
Founder & Senior Journalist • 747 Articles
Barry Elad is a seasoned journalist and analyst specializing in finance, technology, AI, and founder of SQ Magazine. He explores the world o...
LATEST POSTS:
Amazon’s Strands Decider 2B Model Targets Faster AI Agents
Gemini 4 Argon Launched by Google With Cybersecurity Features
Meta Disputes Claim Muse AI Read Private Messages
Robert A. Lee
Senior Editor • 463 Articles
Robert A. Lee is a journalist at SQ Magazine who unpacks the fast-moving worlds of gaming and internet trends. He tracks everything from maj...
LATEST POSTS:
Clash Royale Statistics 2026: Revenue, Players and Engagement
Email Security for Small Businesses: The Five Controls That Matter Most
The Growing Role of Device Fingerprinting in Cybersecurity
What Is a Token in AI

An AI token is the granularity at which Gemini and other generative AI models process input and output. Tokens can be single characters like z or whole words like cat. Long words are broken up into several tokens, according to Google’s Gemini API documentation.

The language-model sense applies throughout this entry: The subword unit a model reads and writes, rather than the blockchain asset that shares the name. The context window defines the combined limit of input and output tokens. When billing is enabled, the cost of a call to the Gemini API is determined in part by the number of input and output tokens.

Key Takeaways

  • For Gemini models, a token is equivalent to about 4 characters, per Google’s Gemini API documentation. 100 tokens is equal to about 60-80 English words.
  • For Claude, a token approximately represents 3.5 English characters, per Anthropic’s documentation glossary. The exact number can vary depending on the language used.
  • Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text, according to Anthropic’s pricing documentation.
  • Knowing how many tokens are in a text string shows whether the string is too long for a text model to process. It also shows how much an OpenAI API call costs, since usage is priced by token.
  • Claude Opus 5 is priced at $5 per million base input tokens and $25 per million output tokens. A cache hit costs 10% of the standard input price.
  • Video is counted at 263 tokens per second and audio at 32 tokens per second. Images with both dimensions less than or equal to 384 pixels count as 258 tokens.

How Does AI Tokenization Work?

The set of all tokens used by the model is called the vocabulary. The process of splitting text into tokens is called tokenization, per Google’s Gemini API documentation. Tokens are the smallest individual units of a language model. They can correspond to words, subwords, characters, or even bytes in the case of Unicode, according to Anthropic’s documentation glossary.

1. Text arrives as a string of characters

Claude is provided with text consisting of a series of characters. That text is encoded into a series of tokens for the model to process. Tokens are typically hidden when interacting with language models at the text level. They become relevant when examining the exact inputs and outputs of a language model.

The step never shows up in a chat window, so the unit surfaces only on an invoice or in a context-limit error.

2. The tokenizer splits it into subword units

Byte pair encoding is a way of converting text into tokens, and it attempts to let the model see common subwords. BPE encodings will often split “encoding” into tokens like “encod” and “ing”, per OpenAI’s tiktoken repository. Peer-reviewed work on neural machine translation introduced the approach, encoding rare and unknown words as sequences of subword units. That segmentation is based on the byte pair encoding compression algorithm.

A tokenizer works like a phrase book of the syllables a language repeats most. A common word costs one entry, and a rare surname costs several. It also works like a postal line that bundles frequent street names in a single pass and spells unfamiliar ones letter by letter.

Newsletter
Don’t chase tech news. We track it for you.

One weekly briefing with the launches, AI developments, and breaches that matter. No filler.

Read by pros at Fortinet, TSMC, Barclays, and Deloitte.

3. Each unit maps to an entry in the vocabulary

Long words are broken up into several tokens. For Gemini models, a token is equivalent to about 4 characters. Byte pair encoding also works on arbitrary text, even text that is not in the tokenizer’s training data.

4. The model reads and generates in those units

Given a text string and an encoding such as cl100k_base, a tokenizer can split the text string into a list of tokens. Each Gemini model has a maximum number of tokens it can handle, and the context window defines the combined limit of input and output tokens.

StepWhat happensWhat the model sees
InputText arrives as a series of charactersNothing yet
TokenizeByte pair encoding splits the text into subword unitsCommon subwords such as “encod” and “ing”
Look upEach unit maps to an entry in the vocabularyVocabulary entries rather than letters
GenerateThe model reads and writes in those unitsOutput tokens counted alongside input tokens

Source: Google Gemini API documentation; OpenAI tiktoken repository; Anthropic Claude documentation glossary.

Tokenization opens the pipeline our large language model explainer walks through end to end.

How Many Tokens Is a Word?

For Gemini models, a token is equivalent to about 4 characters, according to Google’s Gemini API documentation. 100 tokens is equal to about 60-80 English words. Anthropic publishes a rough estimate of its own: 1 token is approximately 4 characters or 0.75 words in English. The exact count varies by language and content type.

Anthropic’s documentation glossary publishes a lower figure again. A token approximately represents 3.5 English characters for Claude, and the exact number can vary depending on the language used. Those ratios come from the vendors themselves, and they disagree. Treating one of them as a constant is where token-cost arithmetic goes wrong.

The spread widens inside a single vendor. Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text. Claude Sonnet 4.6 and earlier models use the previous tokenizer, per Anthropic’s pricing documentation. The exact increase depends on the content and workload shape.

Vendor documentationDocumented characters per tokenDocumented word equivalenceStated qualifier
Google, Gemini API tokens pageAbout 4100 tokens is about 60 to 80 English words“equivalent to about”
Anthropic, Claude documentation glossaryApproximately 3.5Not published“can vary depending on the language used”
Anthropic, pricing documentationApproximately 4Approximately 0.75 words per token“as a rough estimate”

Source: Google Gemini API documentation; Anthropic Claude documentation glossary and pricing documentation.

A price quoted per million tokens is a rate rather than a total. Two sticker prices become comparable only once you know which tokenizer produced the count underneath them.

The two assistants readers ask about most sit on different tokenizers, one gap among several in our Claude and ChatGPT comparison data.

Why Do AI Models Charge Per Token?

Usage is priced by token, according to OpenAI’s Cookbook. Knowing how many tokens are in a text string shows how much an OpenAI API call costs. When billing is enabled, the cost of a call to the Gemini API is determined in part by the number of input and output tokens.

Anthropic denominates its published prices in MTok, which its pricing documentation defines as million tokens.

ModelBase input, per million tokensOutput, per million tokens
Claude Opus 5$5$25
Claude Sonnet 4.6$3$15
Claude Haiku 4.5$1$5

Source: Anthropic Claude pricing documentation, accessed July 2026.

Output carries the heavier rate across the whole line. Claude Opus 5 is priced at $5 per million base input tokens and $25 per million output tokens. Claude Sonnet 4.6 sits at $3 and $15, and Claude Haiku 4.5 at $1 and $5.

Two documented multipliers move the arithmetic the other way. A cache hit costs 10% of the standard input price. The Batch API allows asynchronous processing of large volumes of requests with a 50% discount on both input and output tokens.

Published rates move often, and the count underneath them moves too. Current per-token rates and context sizes across the major models sit in our AI model tracker.

The same unit sets the ceiling as well as the bill. Anthropic’s glossary defines the context window as the amount of text a language model can look back on and reference when generating new text.

Metered per-token billing is the business model underneath the meter, and its revenue side sits in our OpenAI revenue and usage data.

Tokens Beyond Text: Images, Video, and Audio

Images with both dimensions less than or equal to 384 pixels count as 258 tokens. Larger images are tiled into 768×768 pixel tiles, each counting as 258 tokens, according to Google’s Gemini API documentation. Video is counted at 263 tokens per second, and audio is counted at 32 tokens per second.

Input typeDocumented token costScope
Image, both dimensions 384 pixels or less258 tokensGemini API
Image tile, 768×768 pixels258 tokens per tileGemini API
Video263 tokens per secondGemini API
Audio32 tokens per secondGemini API

Source: Google Gemini API documentation, multimodal token counts.

Those are Google’s published rates for the Gemini API, and they do not transfer to other vendors. Non-text inputs also draw on the same budget as text. Gemini and other generative AI models process input and output at a granularity called a token. The context window defines the combined limit of input and output tokens.

Advantages and Trade-offs of Token-Based Accounting

Advantages

  • Byte pair encoding is reversible and lossless, so tokens convert back into the original text.
  • It works on arbitrary text, even text that is not in the tokenizer’s training data.
  • It compresses the text, so the token sequence is shorter than the bytes corresponding to the original text.
  • Larger tokens enable data efficiency during inference and pretraining, and are used when possible.
  • Smaller tokens allow a model to handle uncommon or never-before-seen words.

Trade-offs and Risks

  • The choice of tokenization method can impact the model’s performance and vocabulary size. It also affects the ability to handle out-of-vocabulary words.
  • For Claude, a token approximately represents 3.5 English characters. The exact number can vary depending on the language used. Writers working in some languages pay more units for the same meaning.
  • Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text. Cost baselines built against an older model generation do not carry forward.

Real-World Token Accounting

Anthropic changed the tokenizer between Claude generations

Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. That tokenizer produces approximately 30% more tokens for the same text. Claude Sonnet 4.6 and earlier models use the previous tokenizer.

A team that budgeted against Sonnet 4.6 and then moved up a generation counts different units for an identical prompt.

OpenAI ships a different encoding per model family

Encoding name o200k_base is used by the gpt-4o and gpt-4o-mini models, and cl100k_base by gpt-4-turbo, gpt-4 and gpt-3.5-turbo. Encoding name p50k_base is used by Codex models, and r50k_base by GPT-3 models like davinci.

Counting with the wrong encoding returns a number that looks plausible and bills incorrectly. Adoption across the assistants built on those model families sits in our AI assistant usage and pricing data.

Google publishes a per-second rate for audio and video

Google counts video at 263 tokens per second and audio at 32 tokens per second for the Gemini API. An hour of recorded audio therefore carries a published token cost before anyone writes a prompt around it.

Is an AI Token the Same as a Crypto Token?

No. The AI sense names the granularity at which Gemini and other generative AI models process input and output. That unit counts against the context window and against the cost of a call to the Gemini API.

The blockchain sense describes a digital asset recorded on a ledger. The two senses share a word and nothing else, and no ratio on this page applies to a traded asset.

How Do I Count Tokens Before Sending a Prompt?

tiktoken is a fast open-source tokenizer by OpenAI. Given a text string and an encoding, a tokenizer can split the text string into a list of tokens, according to OpenAI’s Cookbook. Matching the encoding to the model matters. The encoding o200k_base is used by the gpt-4o and gpt-4o-mini models, while cl100k_base is used by gpt-4-turbo, gpt-4, and gpt-3.5-turbo.

A rough sanity check helps before a tokenizer runs. As a rough estimate, 1 token is approximately 4 characters or 0.75 words in English. The exact count varies by language and content type.

Conclusion

An AI token is the unit models read and write in, equivalent to about 4 characters for Gemini models. 100 tokens is equal to about 60-80 English words. For Claude, a token approximately represents 3.5 English characters. Two vendors, two published ratios, and no single conversion that holds across both.

Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text. A per-token price is a rate whose multiplier is set by the model. Budget the tokenizer before budgeting the token.

Definition of AI Inference. Link to full glossary entry follows the description.AI Inference

AI inference is the execution phase where a trained AI model applies what it learned to new, unseen data and produces an output such as a prediction.

Read more

Definition of Context Window. Link to full glossary entry follows the description.Context Window

A context window is all the text an AI model can reference when generating a response, measured in tokens and shared with the model's own output.

Read more

Published on: August 13, 2026

Share ChatGPT Perplexity

Explore More Terms

Context Window

Context Window

A context window is all the text an AI model can reference when generating a response, measured in tokens and shared with the model's own output.

AI Inference

AI Inference

AI inference is the execution phase where a trained AI model applies what it learned to new, unseen data and produces an output such as a prediction.

Multimodal AI

Multimodal AI

Multimodal AI is a single model that takes in and relates more than one type of input, such as text, images, audio or video, rather than only text.

LLM Parameters

LLM Parameters

LLM parameters are the weights and biases a model learns during training. The count sets a model's size, but training data and sparsity matter just as much.

Frontier Model

Frontier Model

A frontier model is a highly capable general-purpose AI model that matches or exceeds today's most advanced systems, and triggers safety obligations.

Mixture of Experts

Mixture of Experts

A mixture of experts (MoE) model splits its feed-forward layers into expert sub-networks and a router sends each token to only a few of them.

Primary Sidebar

Connect With Us

facebook x linkedin google-news telegram pinterest whatsapp email
google-preferred-source-badge Add as a preferred source on Google

You Should Also Read

What Is a Context Window? How AI Models Handle Long Inputs
What Is AI Inference? How a Trained Model Produces Output
What Is Multimodal AI? Text, Vision, and Audio in One Model

Table of Contents

  • Key Takeaways
  • How Does AI Tokenization Work?
  • How Many Tokens Is a Word?
  • Why Do AI Models Charge Per Token?
  • Tokens Beyond Text: Images, Video, and Audio
  • Advantages and Trade-offs of Token-Based Accounting
  • Real-World Token Accounting
  • Is an AI Token the Same as a Crypto Token?
  • How Do I Count Tokens Before Sending a Prompt?
  • Conclusion

Weekly stats quiz Week 41

How much of a tech geek are you?

5 fast questions from this week's verified industry data. About a minute.

Play the quiz New every Monday
Nano Banana 2 1 Launched Ai Search
Artificial Intelligence

Nano Banana 2.1 Lands in Google Search With Big Upgrades

By Barry Elad October 6, 2026
Libreoffice Java Malicious Spreadsheets Flaw Patch
Cybersecurity

LibreOffice Fixes Silent RCE Vulnerability, OpenOffice Still Exposed

By Sofia Ramirez October 6, 2026
Gmo And Mrmax Data Breach
Cybersecurity

GMO Research & AI Breach Hits 948,500 infoQ Members

By Sofia Ramirez October 6, 2026
Apple Autumn Launch Surprise
Technology

Apple Eyes Surprise Late-October Launch for New Macs

By Sofia Ramirez October 6, 2026
Gitlab Cybersecurity Patch Alert
Cybersecurity

GitLab Warns of Critical RCE Flaw in Self-Hosted AI Gateway

By Sofia Ramirez October 2, 2026
Microsoft X Account Hacked Clippy Crypto
Cybersecurity

Microsoft X Hackers Push Unauthorized Clippy Crypto Token

By Sofia Ramirez October 2, 2026
Dell Cybersecurity Infrastructure Alert
Cybersecurity

Dell Patches Two CVSS 10 Container Storage Modules Flaws

By Sofia Ramirez October 2, 2026
Amazon Strands Deciders 2b Ai Model Jev Killer
Artificial Intelligence

Amazon’s Strands Decider 2B Model Targets Faster AI Agents

By Barry Elad October 1, 2026

Footer

SQ Magazine Logo

Smarter Insights for a Fast-Moving Digital World

Connect With Us

Follow Us on Google News

Editorial & Trust

  • About
  • Publishing Principles
  • Fact-Check Policy
  • Corrections Policy
  • Ethics Policy
  • Disclaimer
  • Cookie Policy

Worth Checking

  • The Tech Index
  • The Threat Index
  • Social Media Attention Span Stats
  • Instagram Followers Stats
  • Google Usage Stats
  • LLM Hallucination Stats
  • Gen Z Social Media Stats
Contact Us
13570 Grove Dr #189,
Maple Grove, MN 55311,
United States
10 a.m. to 6 p.m. | Every day

Copyright © 2022–2026 SQ Magazine. All Rights Reserved. Powered by the Neural Stack.

  • Privacy Policy
  • Terms
  • Accessibility Statement
Company
  • About Us
  • Our Team
  • Our Mission
  • Core Values
Discover
  • Brand Assets
    Brand Assets
  • Stats Methodology
    Stats Research Process
  • Glossary
    Glossary
Categories
  • Internet
  • Technology
  • Artificial Intelligence
  • Gaming
  • Cryptocurrency
Internet
Average Attention Span Statistics The Cross-Domain Numbers
Average Attention Span Statistics 2026: The Cross-Domain Numbers
How Many Videos Are on YouTube Statistics
How Many Videos Are on YouTube Statistics 2026: Key Data
How Many People Work at WhatsApp
How Many People Work at WhatsApp 2026: Employee Count and History
Spotify Listening Statistics
Spotify Listening Statistics 2026: Average Listening Time
How Many Subscribers Does MrBeast Have
How Many Subscribers Does MrBeast Have in 2026? Channel Growth Statistics
WhatsApp Business Statistics
WhatsApp Business Statistics 2026: Real Market Insights
Technology
Average US Salary Statistics Median Income by Age and State
Average US Salary Statistics 2026: Median Income by Age and State
Aptoide Statistics 2026: Downloads, Users and App Store Share
Aptoide Statistics 2026: Downloads, Users and App Store Share
AppsFlyer Statistics Customers Revenue and Market Position
AppsFlyer Statistics 2026: Customers, Revenue and Market Position
How Many iPhones Has Apple Sold
How Many iPhones Has Apple Sold in 2026? Units Sold by Year
How Many Employees Does Amazon Have
How Many Employees Does Amazon Have 2026: Workforce Growth
Netflix vs. Hulu Statistics
Netflix vs Hulu Statistics 2026: Viewer Growth Data
Artificial Intelligence
AI Search Engine Statistics Usage Market Share and Adoption
AI Search Engine Statistics 2026: Usage, Market Share and Adoption
AI Music Statistics
AI Music Statistics 2026: Generation, Adoption and Industry Impact
AI Coding Statistics
AI Coding Statistics 2026: Adoption, Productivity and Market Data
How Much Content on Social Media Is AI Generated Statistics
How Much Content on Social Media Is AI Generated Statistics 2026: Hidden Truths
ChatGPT vs DeepSeek Statistics
ChatGPT vs DeepSeek Statistics 2026: Users, Benchmarks & Pricing
ChatGPT vs Claude vs Gemini vs Perplexity Statistics
ChatGPT vs Claude vs Gemini vs Perplexity Statistics 2026: Users, Revenue & Market Share
Gaming
Clash Royale Statistics Revenue Players and Engagement
Clash Royale Statistics 2026: Revenue, Players and Engagement
Gaming Statistics
Gaming Statistics 2026: Market Size, Players, Revenue, and Platforms
Roblox vs Minecraft Statistics
Roblox vs Minecraft Statistics 2026: Players, Revenue, Creators
Online Gambling Regulations Statistics
Online Gambling Regulations Statistics 2026: Global Compliance and Enforcement Data
Fantasy Sports Statistics
Fantasy Sports Statistics 2026: Users, Revenue & Trends
Apex Legends Statistics 2026: Players, Revenue, and Esports
Apex Legends Statistics 2026: Players, Revenue, and Esports
Cryptocurrency
How Many Bitcoins Are There
How Many Bitcoins Are There in 2026? Supply, Mined and Remaining Statistics
Stablecoin Usage Statistics
Stablecoin Usage Statistics 2026: Explosive Growth
Cryptocurrency Adoption Statistics
Cryptocurrency Adoption Statistics 2026: Shocking Trends Now
Coinbase Wallet Statistics
Coinbase Wallet Statistics 2026: Users, Security
Dogecoin Statistics 2026: Circulating Supply, Inflation Rate and ETFs
Dogecoin Statistics 2026: Circulating Supply, Inflation Rate and ETFs
BONK Coin Statistics Supply Holders Price and Treasury Data
BONK Coin Statistics 2026: Supply, Holders, Price and Treasury Data
Categories
  • Artificial Intelligence
  • Cybersecurity
  • Technology
  • Internet
  • Cryptocurrency
Artificial Intelligence
Nano Banana 2 1 Launched Ai Search
Nano Banana 2.1 Lands in Google Search With Big Upgrades
Amazon Strands Deciders 2b Ai Model Jev Killer
Amazon’s Strands Decider 2B Model Targets Faster AI Agents
Gemini 4 Argon Launch
Gemini 4 Argon Launched by Google With Cybersecurity Features
Meta Dispute Muse Private Message Read Claim
Meta Disputes Claim Muse AI Read Private Messages
Openai Dots Gpt 6 1 Sol Launched
OpenAI Launches Dots Agents and GPT-6.1 Sol at DevDay 2026
Claude Down Signin Issues
Claude Down: Anthropic Outage Disrupts Chats, Code and Sign-Ins
Cybersecurity
Libreoffice Java Malicious Spreadsheets Flaw Patch
LibreOffice Fixes Silent RCE Vulnerability, OpenOffice Still Exposed
Gmo And Mrmax Data Breach
GMO Research & AI Breach Hits 948,500 infoQ Members
Gitlab Cybersecurity Patch Alert
GitLab Warns of Critical RCE Flaw in Self-Hosted AI Gateway
Microsoft X Account Hacked Clippy Crypto
Microsoft X Hackers Push Unauthorized Clippy Crypto Token
Dell Cybersecurity Infrastructure Alert
Dell Patches Two CVSS 10 Container Storage Modules Flaws
Europol Cybercrime Takedown
Europol Targets KillSec in Massive Ransomware Operation
Technology
Apple Autumn Launch Surprise
Apple Eyes Surprise Late-October Launch for New Macs
Microsoft Windows Deployment Service Deprecation
Microsoft Will Deprecate Windows Deployment Services After Server 2025
Youtube Custom Feeds With Gemini Ai
YouTube’s AI Feed Builder Changes Video Discovery
Iphone 18 Pro Face Id Bug Reboot Crash
New iPhone 18 Pro Bug Makes Face ID Crash and Reboot
Googlebook With Gemini Ai Launched
Googlebook’s Bold Laptop Launch Starts at $899 in the US
New Samsung Patent Reveals Galaxy Watch Glucose Tracking
New Samsung Patent Reveals Galaxy Watch Glucose Tracking
Internet
Meta Launched Meta One Subscription
Meta One Bundles Instagram, Facebook, WhatsApp Into One AI Subscription
Apple Wallet Ids Launch In Oklahoma
Apple Wallet IDs Launch in Oklahoma in Major Expansion
Meta to Pay 18 Billion in Landmark Teen Safety Deal
Meta to Pay $18 Billion in Landmark Teen Safety Deal
Whatsapp Brings Passkeys 2fa
WhatsApp Hits 1 Billion Passkey Users, Adds 2FA Passwords
Apple Eu App Store Fee Reduction
Apple Sets New EU App Store Fees, Effective October 1
Github Outage Aug 2026
GitHub Down: Outage Hits Thousands of Users Worldwide
Cryptocurrency
Sonic Labs Launch Ussd Stablecoin
Sonic Launches USSD Stablecoin Backed by US Treasuries
Bhutan Moves 12m In Bitcoins
Bhutan Moves $12 Million in Bitcoin from Primary Wallets
Curve Finance Accuses Pancakeswap For Code Stealing
Curve Accuses PancakeSwap of Copying StableSwap Code
Strike Receives Bitlicense In New York
Strike Gets New York BitLicense for Bitcoin Financial Services
Scotiabank Multi Crypto Etf 3iqlogos
Scotiabank Launches Multi Crypto ETF With 3iQ in Canada
Nyse Parent Invests In Okx Crypto Exchange
ICE Invests in OKX to Bridge Crypto and Traditional Finance
Newsletter

Too much tech noise?

We respect your time. One high-signal briefing a week: tech, AI, and security. Nothing else.

Read by pros at Fortinet, TSMC, Barclays, and Deloitte.
Newsletter

The SQ Briefing

We track tech, AI, and security 24/7. You get a 5-minute weekly summary.

Read by pros at Fortinet, TSMC, Barclays, and Deloitte.