SQ Magazine The Threat Index Tracker

AI Model Tracker

A curated timeline of the models that matter: who shipped what, the specs and license you can actually build on, and what it costs. Every entry is sourced from the developer's own release.

19 models tracked verified through 25 Jul 2026 next review by 24 Aug 2026

Latest change Claude Opus 5 added in release week: $5/$25 pricing, 1M context, May 2026 knowledge cutoff, developer-reported benchmarks from the system card. Claude Opus 4.8 marked superseded but still available. (24 Jul 2026) All changes

Models tracked
19
Open weights
3
Developers
10
Re-verify cycle
30day
A sourced timeline of model releases, not a benchmark board. Specs and pricing are as published by the developer; “undisclosed” stays undisclosed, and price is a relative tier with the exact figure on the provider’s page via each row’s source. Informational only.
Model Developer Released Modality License Price Access Our coverage Record detail
GPT-5.6 Sol OpenAI Jul 2026 Text + Vision Proprietary $$$ API OpenAI Launches GPT 5.6 Sol With Powerful New AI Features →
Context window
1.1M tokens
Parameters
Undisclosed
Knowledge cutoff
Feb 2026
Best at
Flagship of the GPT-5.6 series (Sol/Terra/Luna)
Benchmarks (developer-reported)
SWE-Bench Pro 64.6% · GPQA Diamond 94.6% · Terminal-Bench 2.1 88.8%
License terms
Commercial API; OpenAI terms
Primary source
OpenAI release ↗
Last confirmed
24 Jul 2026
Recent change
Row represents GPT-5.6 Sol, the frontier variant; the gpt-5.6 alias routes to Sol. Series also includes Terra (mid, 2.50/15) and Luna (cost-optimized, 1/6). All 1.05M context, Feb 2026 cutoff.
Security & safety
System card and preparedness framework; enterprise data controls.
Gemini 3.6 Flash Google Jul 2026 Multimodal Proprietary $ API Google model card ↗
Context window
1M tokens
Parameters
Undisclosed
Knowledge cutoff
Mar 2026
Best at
Fast workhorse, up to 17% fewer tokens
Benchmarks (developer-reported)
SWE-Bench Pro 58.7% · OSWorld-Verified 83.0% · Terminal-Bench 2.1 78.0%
License terms
Commercial API; Google terms
Primary source
Google model card ↗
Last confirmed
24 Jul 2026
Recent change
Builds on Gemini 3.5 Flash; ~17% fewer output tokens. Model card published Jul 21 2026. Gemini 3.5 Pro in partner testing.
Security & safety
Model card and safety evals; enterprise data controls.
Gemini 3.5 Flash-Lite Google Jul 2026 Multimodal Proprietary $ API Google model card ↗
Context window
1M tokens
Parameters
Undisclosed
Knowledge cutoff
Mar 2026
Best at
Cheapest, low-latency, high throughput
Benchmarks (developer-reported)
SWE-Bench Pro 54.2% · OSWorld-Verified 74.0% · Terminal-Bench 2.1 54.0%
License terms
Commercial API
Primary source
Google model card ↗
Last confirmed
24 Jul 2026
Recent change
Based on Gemini 3.1 Flash-Lite; fastest 3.5-class model (350 tokens/s), up to 1M context per model card.
Security & safety
Model card; safety evals.
Claude Opus 5 Anthropic Jul 2026 Text + Vision Proprietary $$$ API Anthropic release ↗
Context window
1M tokens
Parameters
Undisclosed
Knowledge cutoff
May 2026
Best at
Complex agentic coding and enterprise work
Benchmarks (developer-reported)
SWE-bench Verified 96.0% · OSWorld 2.0 70.6% · DeepSWE v1.1 68.8%
License terms
Commercial API
Primary source
Anthropic release ↗
Last confirmed
25 Jul 2026
Recent change
Added in release week; $5/$25 pricing, 1M context, May 2026 knowledge cutoff per the models overview. SWE-bench Verified figure appears in system card section 8.2 prose, not the summary table
Security & safety
System card published. Cyber classifiers allow source-code vulnerability discovery but block binary scanning, penetration testing, and exploit generation.
Claude Fable 5 Anthropic Jun 2026 Text + Vision Proprietary $$$ API Claude Fable 5 Ends Free Access For Pro Subscribers →
Context window
1M tokens
Parameters
Undisclosed
Knowledge cutoff
Jan 2026
Best at
Frontier intelligence for long-running agents
License terms
Commercial API; Mythos-class 30-day data retention, not used for training
Primary source
Anthropic release ↗
Last confirmed
24 Jul 2026
Security & safety
System card published; enterprise data controls.
Claude Sonnet 5 Anthropic Jun 2026 Text + Vision Proprietary $$ API Anthropic release ↗
Context window
1M tokens
Parameters
Undisclosed
Knowledge cutoff
Jan 2026
Best at
Most agentic Sonnet; near-Opus at lower cost
Benchmarks (developer-reported)
SWE-bench Verified 85.2% · OSWorld-Verified 81.2% · Terminal-Bench 2.1 80.4%
License terms
Commercial API; Anthropic usage policy
Primary source
Anthropic release ↗
Last confirmed
24 Jul 2026
Recent change
Upgrade to Sonnet 4.6; standard pricing 3/15, introductory 2/10 through Aug 31 2026. 1M context, 128K output. Released Jun 30 2026.
Security & safety
System card published; cyber safeguards enabled by default; enterprise data controls.
Claude Opus 4.8 Anthropic May 2026 Text + Vision Proprietary $$$ API Anthropic Unveils Claude Science to Transform Research →
Context window
1M tokens
Parameters
Undisclosed
Knowledge cutoff
Jan 2026
Best at
Frontier reasoning, long-horizon agents
Benchmarks (developer-reported)
SWE-bench Verified 88.6% · GPQA Diamond 93.6% · OSWorld-Verified 83.4%
License terms
Commercial API; Anthropic usage policy
Last confirmed
24 Jul 2026
Recent change
Superseded by Claude Opus 5 (Jul 2026); remains available and priced at $5/$25 on the current models page
Security & safety
System card and Responsible Scaling Policy published; customer API data not used for training.
Gemini 3 Pro Google Nov 2025 Multimodal Proprietary API Google model card ↗
Context window
1M tokens
Parameters
Undisclosed
Best at
Frontier multimodal reasoning with Deep Think
Benchmarks (developer-reported)
SWE-bench Verified 76.2% · GPQA Diamond 91.9% · Terminal-Bench 2.0 54.2%
License terms
Commercial API
Primary source
Google model card ↗
Last confirmed
24 Jul 2026
Recent change
Base of the Gemini 3 Pro family. Not on the current Google Gemini API pricing page (superseded by Gemini 3.1 Pro at 2/12 and Gemini 3.5); no longer first-party purchasable, so price cleared under the current-pricing test.
Security & safety
Model card and safety evals; enterprise data controls.
Claude Haiku 4.5 Anthropic Oct 2025 Text + Vision Proprietary $ API Anthropic model card ↗
Context window
200K tokens
Parameters
Undisclosed
Knowledge cutoff
Feb 2025
Best at
Low-latency, cost-efficient
Benchmarks (developer-reported)
SWE-bench Verified 73.3% · GPQA Diamond 73.0% · AIME 2025 80.7% (no tools)
License terms
Commercial API
Last confirmed
24 Jul 2026
Security & safety
System card; enterprise data controls.
GPT-5 OpenAI Aug 2025 Text + Vision Proprietary $$ API OpenAI Launches ChatGPT Work With GPT-5.6 Agents →
Context window
400K tokens
Parameters
Undisclosed
Knowledge cutoff
Sep 2024
Best at
Broad reasoning and tools
Benchmarks (developer-reported)
SWE-bench Verified 74.9% · AIME 2025 94.6% (no tools) · MMMU 84.2%
License terms
Commercial API
Primary source
OpenAI model card ↗
Last confirmed
24 Jul 2026
Recent change
Superseded by GPT-5.6 (Jul 2026); OpenAI lists GPT-5 as prior-generation.
Security & safety
System card published.
Grok 2 xAI Aug 2025 Text Community Self-host Open weights xAI (Hugging Face) ↗
Context window
128K tokens
Parameters
Undisclosed
Best at
Open frontier-class weights
Benchmarks (developer-reported)
HumanEval 88.4% · GPQA 56.0% · MMLU-Pro 75.5%
License terms
Grok 2 Community License Agreement
Last confirmed
24 Jul 2026
Recent change
Open-weights release Aug 2025 (HF repo xai-org/grok-2 created 2025-08-22). Context 128K from repo config.json (max_position_embeddings 131072, RoPE-scaled from 8192). Sparse MoE (8 experts, 2 active per token); xAI publishes no headline total-parameter count, so params left absent. Text-only causal LM (Grok1ForCausalLM). Non-commercial Grok 2 Community License; weights-only, no first-party API.
Security & safety
Open weights under the Grok 2 Community License (restricted); deployer owns safety tuning.
Llama 4 Meta Apr 2025 Text + Vision Community Self-host Open weights Meta model card ↗
Context window
1M tokens
Parameters
up to 400B (MoE)
Best at
Open MoE, long context
Benchmarks (developer-reported)
LiveCodeBench 43.4% · GPQA Diamond 69.8% · MMLU-Pro 80.5%
License terms
Llama Community License (scale cap)
Primary source
Meta model card ↗
Last confirmed
24 Jul 2026
Recent change
Meta operates no first-party inference API and points developers to third-party hosts, so access is weights-only. Represents Llama 4 Maverick (1M context, 17B active / 400B total, 128 experts, released 2025-04-05); Scout variant up to 10M context. Text plus image, no audio or video.
Security & safety
Open weights: you own the data boundary and the safety tuning. Scale-cap clause for very large deployments.
Qwen3 Alibaba Apr 2025 Text Open weights Self-host Open weights Alibaba model card ↗
Context window
128K tokens
Parameters
0.6B-235B
Best at
Broad open family, multilingual
Benchmarks (developer-reported)
LiveCodeBench 70.7% · AIME 2025 81.5% · BFCL v3 70.8%
License terms
Apache-2.0
Last confirmed
24 Jul 2026
Recent change
Alibaba current Model Studio API lists qwen3.7-max/plus and qwen3.6-flash; the open Apr-2025 variant qwen3-235b-a22b is not on the current first-party pricing page, so price kept free and access set to weights-only (Apache-2.0). Flagship Qwen3-235B-A22B at 128K context; base Qwen3 is text (Qwen3-VL/Omni add vision). Later 2507 updates raised context to 256K (1M on select variants).
Security & safety
Apache-2.0 weights; self-host data boundary.
Gemma 3 Google Mar 2025 Text + Vision Community Self-host Open weights Google model card ↗
Context window
128K tokens
Parameters
1B-27B
Best at
Efficient open family
Benchmarks (developer-reported)
LiveCodeBench 29.7% · GPQA Diamond 42.4% · MMLU-Pro 67.5%
License terms
Gemma terms of use (restricted)
Primary source
Google model card ↗
Last confirmed
24 Jul 2026
Recent change
Open family (1B/4B/12B/27B); 4B/12B/27B support vision (text plus image). Built on Gemini 2.0 research. ShieldGemma 2 image safety checker built on Gemma 3.
Security & safety
Open weights under Gemma terms; self-host safety is your responsibility.
DeepSeek-R1 DeepSeek Jan 2025 Text Open weights $ API + weights DeepSeek model card ↗
Context window
128K tokens
Parameters
671B (MoE)
Best at
Open reasoning model
Benchmarks (developer-reported)
SWE-bench Verified 49.2% · GPQA Diamond 71.5% · MMLU-Pro 84.0%
License terms
MIT
Last confirmed
24 Jul 2026
Recent change
Updated as R1-0528 (May 2025).
Security & safety
MIT weights; self-host data boundary. Reasoning traces can be verbose; review before logging.
DeepSeek-V3 DeepSeek Dec 2024 Text Community $ API + weights DeepSeek model card ↗
Context window
128K tokens
Parameters
671B (MoE)
Best at
Efficient open MoE, low-cost API
Benchmarks (developer-reported)
SWE-bench Verified 42.0% · GPQA Diamond 59.1% · MMLU-Pro 75.9%
License terms
DeepSeek Model License (code is MIT; weights carry use restrictions)
Last confirmed
24 Jul 2026
Recent change
License corrected: V3 weights ship under the DeepSeek Model License, not MIT — MIT covers only the code repository (R3 audit, Jul 2026)
Security & safety
MIT weights plus low-cost hosted API. Self-host for data control; review provider data policy if using the API.
Phi-4 Microsoft Dec 2024 Text Open weights Self-host Open weights Microsoft model card ↗
Context window
16K tokens
Parameters
14B
Knowledge cutoff
Jun 2024
Best at
Small-model reasoning
Benchmarks (developer-reported)
HumanEval 82.6% · GPQA 56.1% · MMLU 84.8%
License terms
MIT
Last confirmed
24 Jul 2026
Recent change
Phi-4 family expanded in 2025 (Phi-4-mini, Phi-4-multimodal, Phi-4-reasoning).
Security & safety
MIT weights, small enough to run locally; a good fit where data cannot leave the device.
Command R+ Cohere Aug 2024 Text Community $$ API + weights Cohere model card ↗
Context window
128K tokens
Parameters
104B
Best at
RAG and tool use
License terms
CC-BY-NC (non-commercial)
Primary source
Cohere model card ↗
Last confirmed
24 Jul 2026
Recent change
Command R+ 08-2024 (current version), a refresh of the original Apr 2024 release; 128K context, CC-BY-NC weights plus paid Cohere API. Superseded by Command A (Mar 2025) and Command A Reasoning.
Security & safety
Weights non-commercial (CC-BY-NC); enterprise license required to ship commercially.
Mistral Large 2 Mistral Jul 2024 Text Community Self-host Open weights Mistral documentation ↗
Context window
128K tokens
Parameters
123B
Best at
Strong European frontier
Benchmarks (developer-reported)
MMLU 84.0% (pretrained)
License terms
Mistral Research License (non-commercial); commercial license to self-deploy
Last confirmed
24 Jul 2026
Recent change
First-party API mistral-large-2407 deprecated 2024-11-30 and retired 2025-03-30 per Mistral docs; superseded by Mistral Large 3 (25.12). Weights remain under Mistral Research License (non-commercial); price cleared and access set to weights-only.
Security & safety
EU-based provider; commercial license with enterprise data terms.

Verification ledger

4 most recent of 4 logged updates
  • Claude Opus 5 added in release week: $5/$25 pricing, 1M context, May 2026 knowledge cutoff, developer-reported benchmarks from the system card. Claude Opus 4.8 marked superseded but still available. 24 Jul 2026
  • Added developer-reported release benchmarks to 17 model records; two models (Claude Fable 5, Command R+) published no standard benchmark table and stay without one. 24 Jul 2026
  • Round-3 audit correction: DeepSeek-V3 reclassified from open weights to community license — its weights ship under the DeepSeek Model License; MIT covers only the code repository. 24 Jul 2026
  • Tracker launched: 18 model records published after a three-round verification against developer sources; 2 records held pending public documentation. 24 Jul 2026

How this tracker is maintained

Every model passes the same checks before it appears, and the row keeps pace with each new release.

  1. Sourced

    Specs, license and pricing come from the developer’s own release: model card, license file, or pricing page. No benchmark screenshots, no third-hand numbers.

  2. Dated

    Each row carries its release date and when we last confirmed it. Each version bump gets its own new row.

  3. Re-checked

    Reviewed every 30 days because the field moves fast. License changes, price cuts and deprecations are logged in the ledger.

Why not a leaderboard?
Live benchmarks already exist and shift daily. This is the durable record: what shipped, when, under which license, at what cost. Row details carry the scores each developer published at release; for current head-to-head rankings, use a live board such as LMArena.
How often is it updated?
Every 30 days, and immediately when a major model ships. The ledger logs each addition and reclassification.
Can I cite this?
Yes. Every row links the developer’s primary source. Use “Cite this tracker” for a ready reference.

Informational only. Specs and pricing reflect developer disclosures at publish time and can change with new versions.

Sources