Skip to content

Intelligence brief · Evidence current to 29 July 2026

AI is moving from models to systems.

The newest releases combine models, tools and distribution. At the same time, regulation is becoming operational—and Nepal is assembling the policy, talent and language resources needed to participate on its own terms.

Evidence standard

6 of 13 editorial records are cross-checked, primary-source verified or peer-reviewed. Commercial performance claims stay labelled as provider claims.

Last research verification · 29 July 2026

01 · Major current findings

What changed—and why it matters now

Findings are derived from the curated records displayed in this application, not from generic market commentary.

01

2 Aug 2026

Governance becomes operational

EU transparency duties begin days after the AI Omnibus entered into force.

02

4 of 5

Agentic becomes the default release story

Four of the five newest tracked model releases foreground tools, actions, agents or computer use.

03

3 layers

Nepal has a visible foundation

The validated record now connects national policy, dedicated AI education and Nepali-language model research.

04

0 launch claims

Independent evaluation is the bottleneck

None of the newest commercial model launch claims in this curated set is treated as independently reproduced.

02 · Model intelligence over time

Twelve charts on how models, prices and labs are actually moving.

Hover or keyboard-focus any point or bar for exact values, model names and dates.

Release cadence by kind

Model releases per year, split by frontier, open-weight, small and multimodal.

01

2019: {"year":2019,"Frontier":0,"Open-weight":1,"Small / on-device":0,"Multimodal":0}, 2020: {"year":2020,"Frontier":1,"Open-weight":0,"Small / on-device":0,"Multimodal":0}, 2022: {"year":2022,"Frontier":0,"Open-weight":2,"Small / on-device":0,"Multimodal":0}, 2023: {"year":2023,"Frontier":12,"Open-weight":15,"Small / on-device":1,"Multimodal":1}, 2024: {"year":2024,"Frontier":23,"Open-weight":67,"Small / on-device":2,"Multimodal":15}, 2025: {"year":2025,"Frontier":15,"Open-weight":10,"Small / on-device":0,"Multimodal":3}

Model kind mix

Share of all tracked models by kind.

02
100percent
  • Open-weight56.5
  • Frontier30.4
  • Multimodal11.3
  • Small / on-device1.8

Open-weight: 56.5%, Frontier: 30.4%, Multimodal: 11.3%, Small / on-device: 1.8%

Context window growth

Context window (thousands of tokens, log scale) for every model with a disclosed limit.

03

GPT-NeoX-20B (2022): 2K, BLOOM-176B (2022): 2K, GPT-3.5 Turbo (2023): 4K, Jurassic-2 Ultra (2023): 8K, GPT-4 (2023): 8K, GPT-4 32K (2023): 32K, StarCoder (2023): 8K, MPT-7B (2023): 2K, PaLM 2 (2023): 8K, Falcon-40B (2023): 2K, Falcon-7B (2023): 2K, MPT-30B (2023): 8K, Claude 2 (2023): 100K, Llama 2 13B (2023): 4K, Llama 2 70B (2023): 4K, Llama 2 7B (2023): 4K, Falcon-180B (2023): 2K, Mistral 7B (2023): 8K, ChatGLM3-6B (2023): 8K, Yi-34B (2023): 4K, Yi-6B (2023): 4K, GPT-4 Turbo (2023): 128K, Claude 2.1 (2023): 200K, Titan Text Express (2023): 8K, Titan Text Lite (2023): 4K, Gemini 1.0 Pro (2023): 32K, Mixtral 8x7B (2023): 32K, Phi-2 (2023): 2K, GLM-4 (2024): 128K, TinyLlama-1.1B (2024): 2K, StableLM 2 1.6B (2024): 4K, Qwen1.5-0.5B (2024): 32K, Qwen1.5-14B (2024): 32K, Qwen1.5-72B (2024): 32K, Qwen1.5-7B (2024): 32K, Gemma 2B (2024): 8K, Gemma 7B (2024): 8K, Mistral Large (2024): 32K, StarCoder2-15B (2024): 16K, Claude 3 Opus (2024): 200K, Claude 3 Sonnet (2024): 200K, Inflection-2.5 (2024): 32K, Command R (2024): 128K, Claude 3 Haiku (2024): 200K, DBRX (2024): 32K, Grok-1.5 (2024): 128K, Qwen1.5-110B (2024): 32K, Qwen1.5-32B (2024): 32K, Command R+ (2024): 128K, Reka Core (2024): 128K, Reka Flash (2024): 128K, Mixtral 8x22B (2024): 64K, Llama 3 70B (2024): 8K, Llama 3 8B (2024): 8K, Phi-3-medium (2024): 128K, Phi-3-mini (2024): 4K, Phi-3-small (2024): 8K, Snowflake Arctic (2024): 4K, Yi-1.5-34B (2024): 4K, DeepSeek-V2 (2024): 128K, Granite Code 34B (2024): 8K, Granite Code 8B (2024): 128K, GPT-4o (2024): 128K, Falcon2-11B (2024): 8K, Gemini 1.5 Flash (2024): 1000K, Gemini 1.5 Pro (2024): 1000K, Aya 23 35B (2024): 8K, Aya 23 8B (2024): 8K, Codestral (2024): 32K, DeepSeek-Coder-V2 (2024): 128K, GLM-4-9B (2024): 128K, Qwen2-0.5B (2024): 32K, Qwen2-57B-A14B (2024): 64K, Qwen2-72B (2024): 128K, Qwen2-7B (2024): 128K, Nemotron-4 340B (2024): 4K, Claude 3.5 Sonnet (2024): 200K, Gemma 2 27B (2024): 8K, Gemma 2 9B (2024): 8K, Mistral NeMo (2024): 128K, GPT-4o mini (2024): 128K, Minitron-8B (2024): 4K, Llama 3.1 405B (2024): 128K, Llama 3.1 70B (2024): 128K, Llama 3.1 8B (2024): 128K, Mistral Large 2 (2024): 128K, Gemma 2 2B (2024): 8K, Jamba 1.5 Large (2024): 256K, Jamba 1.5 Mini (2024): 256K, GLM-4-Plus (2024): 128K, Grok-2 (2024): 128K, Grok-2 mini (2024): 128K, Phi-3.5-mini (2024): 128K, Qwen2-VL-7B (2024): 32K, DeepSeek-V2.5 (2024): 128K, o1-mini (2024): 128K, o1-preview (2024): 128K, Pixtral 12B (2024): 128K, Qwen2.5-0.5B (2024): 32K, Qwen2.5-1.5B (2024): 32K, Qwen2.5-14B (2024): 128K, Qwen2.5-32B (2024): 128K, Qwen2.5-3B (2024): 32K, Qwen2.5-72B (2024): 128K, Qwen2.5-7B (2024): 128K, Qwen2.5-Coder-7B (2024): 128K, Qwen2-VL-72B (2024): 32K, Llama 3.2 11B Vision (2024): 128K, Llama 3.2 1B (2024): 128K, Llama 3.2 3B (2024): 128K, Llama 3.2 90B Vision (2024): 128K, Gemini 1.5 Flash-8B (2024): 1000K, Ministral 3B (2024): 128K, Ministral 8B (2024): 128K, Granite 3.0 2B (2024): 128K, Granite 3.0 8B (2024): 128K, Claude 3.5 Sonnet (Oct 2024) (2024): 200K, SmolLM2-1.7B (2024): 8K, Claude 3.5 Haiku (2024): 200K, Qwen2.5-Coder-32B (2024): 128K, QwQ-32B-Preview (2024): 32K, Command R7B (2024): 128K, Amazon Nova Lite (2024): 300K, Amazon Nova Micro (2024): 128K, Amazon Nova Pro (2024): 300K, o1 (2024): 200K, Llama 3.3 70B (2024): 128K, Phi-4 (2024): 16K, DeepSeek-V3 (2024): 128K, DeepSeek-R1 (2025): 128K, Mistral Small 3 (2025): 32K, o3-mini (2025): 200K, Gemini 2.0 Flash (2025): 1000K, Gemini 2.0 Flash-Lite (2025): 1000K, Grok 3 (2025): 128K, Grok 3 mini (2025): 128K, Claude 3.7 Sonnet (2025): 200K, GPT-4.5 preview (2025): 128K, Jamba Mini 1.6 (2025): 256K, Command A (2025): 256K, Llama 4 Maverick (2025): 1000K, Llama 4 Scout (2025): 10000K, GPT-4.1 (2025): 1000K, GPT-4.1 (2025): 1000K, GPT-4.1 mini (2025): 1000K, GPT-4.1 nano (2025): 1000K, o3 (2025): 200K, o4-mini (2025): 200K, Amazon Nova Premier (2025): 1000K, Claude Opus 4 (2025): 200K, Claude Sonnet 4 (2025): 200K

Parameter scale growth

Parameter count (billions, log scale) by release year.

04

GPT-2 (2019): 1.5B, GPT-3 (2020): 175B, GPT-NeoX-20B (2022): 20B, BLOOM-176B (2022): 176B, StarCoder (2023): 15.5B, MPT-7B (2023): 7B, Falcon-40B (2023): 40B, Falcon-7B (2023): 7B, MPT-30B (2023): 30B, Llama 2 13B (2023): 13B, Llama 2 70B (2023): 70B, Llama 2 7B (2023): 7B, Falcon-180B (2023): 180B, Mistral 7B (2023): 7B, ChatGLM3-6B (2023): 6B, Yi-34B (2023): 34B, Yi-6B (2023): 6B, Gemini 1.0 Nano (2023): 3.2B, Mixtral 8x7B (2023): 46.7B, Phi-2 (2023): 2.7B, TinyLlama-1.1B (2024): 1.1B, StableLM 2 1.6B (2024): 1.6B, Qwen1.5-0.5B (2024): 0.5B, Qwen1.5-14B (2024): 14B, Qwen1.5-72B (2024): 72B, Qwen1.5-7B (2024): 7B, Gemma 2B (2024): 2B, Gemma 7B (2024): 7B, StarCoder2-15B (2024): 15B, Grok-1 (2024): 314B, DBRX (2024): 132B, Qwen1.5-110B (2024): 110B, Qwen1.5-32B (2024): 32B, Command R+ (2024): 104B, Mixtral 8x22B (2024): 141B, Llama 3 70B (2024): 70B, Llama 3 8B (2024): 8B, Phi-3-medium (2024): 14B, Phi-3-mini (2024): 3.8B, Phi-3-small (2024): 7B, OpenELM-270M (2024): 0.3B, OpenELM-3B (2024): 3B, Snowflake Arctic (2024): 480B, Yi-1.5-34B (2024): 34B, DeepSeek-V2 (2024): 236B, Granite Code 34B (2024): 34B, Granite Code 8B (2024): 8B, Falcon2-11B (2024): 11B, Aya 23 35B (2024): 35B, Aya 23 8B (2024): 8B, Codestral (2024): 22B, DeepSeek-Coder-V2 (2024): 236B, GLM-4-9B (2024): 9B, Qwen2-0.5B (2024): 0.5B, Qwen2-57B-A14B (2024): 57B, Qwen2-72B (2024): 72B, Qwen2-7B (2024): 7B, Nemotron-4 340B (2024): 340B, Gemma 2 27B (2024): 27B, Gemma 2 9B (2024): 9B, Mistral NeMo (2024): 12B, Minitron-8B (2024): 8B, Llama 3.1 405B (2024): 405B, Llama 3.1 70B (2024): 70B, Llama 3.1 8B (2024): 8B, Mistral Large 2 (2024): 123B, Gemma 2 2B (2024): 2B, Jamba 1.5 Large (2024): 398B, Jamba 1.5 Mini (2024): 52B, Phi-3.5-mini (2024): 3.8B, Qwen2-VL-7B (2024): 7B, DeepSeek-V2.5 (2024): 236B, Pixtral 12B (2024): 12B, Qwen2.5-0.5B (2024): 0.5B, Qwen2.5-1.5B (2024): 1.5B, Qwen2.5-14B (2024): 14B, Qwen2.5-32B (2024): 32B, Qwen2.5-3B (2024): 3B, Qwen2.5-72B (2024): 72B, Qwen2.5-7B (2024): 7B, Qwen2.5-Coder-7B (2024): 7B, Qwen2-VL-72B (2024): 72B, Llama 3.2 11B Vision (2024): 11B, Llama 3.2 1B (2024): 1B, Llama 3.2 3B (2024): 3B, Llama 3.2 90B Vision (2024): 90B, Ministral 3B (2024): 3B, Ministral 8B (2024): 8B, Granite 3.0 2B (2024): 2B, Granite 3.0 8B (2024): 8B, SmolLM2-1.7B (2024): 1.7B, Qwen2.5-Coder-32B (2024): 32B, QwQ-32B-Preview (2024): 32B, Command R7B (2024): 7B, Llama 3.3 70B (2024): 70B, Phi-4 (2024): 14B, DeepSeek-V3 (2024): 671B, DeepSeek-R1 (2025): 671B, DeepSeek-R1-Distill-Llama-70B (2025): 70B, DeepSeek-R1-Distill-Llama-8B (2025): 8B, DeepSeek-R1-Distill-Qwen-1.5B (2025): 1.5B, DeepSeek-R1-Distill-Qwen-14B (2025): 14B, DeepSeek-R1-Distill-Qwen-32B (2025): 32B, DeepSeek-R1-Distill-Qwen-7B (2025): 7B, Mistral Small 3 (2025): 24B, Jamba Mini 1.6 (2025): 52B, Command A (2025): 111B, Llama 4 Maverick (2025): 400B, Llama 4 Scout (2025): 109B

Benchmark capability over time

Average public benchmark score by year, for benchmarks measured across multiple release years.

05

GPQA Diamond 2024: 59%, GPQA Diamond 2025: 71%, GSM8K 2023: 93%, GSM8K 2024: 91%, HumanEval 2023: 77%, HumanEval 2024: 82%, HumanEval 2025: 97%, MATH 2024: 87%, MATH 2025: 96%, MMLU 2023: 78%, MMLU 2024: 83%, MMLU 2025: 85%, SWE-bench Verified 2024: 49%, SWE-bench Verified 2025: 66%

Open-weight share of releases

Share of each year's dated releases shipped with open weights.

06

2019: 100% of 1, 2020: 0% of 1, 2022: 100% of 2, 2023: 51.7% of 29, 2024: 62.6% of 107, 2025: 35.7% of 28

Token price over time

Disclosed input and output price per million tokens (USD, log scale), by release year.

07

GPT-3.5 Turbo (2023): in $1.5, out $2, GPT-4 (2023): in $30, out $60, GPT-4 32K (2023): in $60, out $120, GPT-4 Turbo (2023): in $10, out $30, Amazon Nova Lite (2024): in $0.06, out $0.24, Amazon Nova Micro (2024): in $0.035, out $0.14, Amazon Nova Pro (2024): in $0.8, out $3.2, Claude 3.5 Haiku (2024): in $0.8, out $4, Claude 3.5 Sonnet (Oct 2024) (2024): in $3, out $15, Claude 3 Haiku (2024): in $0.25, out $1.25, Claude 3 Opus (2024): in $15, out $75, Claude 3 Sonnet (2024): in $3, out $15, Gemini 1.5 Flash (2024): in $0.075, out $0.3, Gemini 1.5 Flash-8B (2024): in $0.0375, out $0.15, Gemini 1.5 Pro (2024): in $3.5, out $10.5, Mistral Large (2024): in $8, out $24, GPT-4o (2024): in $5, out $15, GPT-4o mini (2024): in $0.15, out $0.6, o1 (2024): in $15, out $60, o1-mini (2024): in $3, out $12, o1-preview (2024): in $15, out $60, Claude 3.7 Sonnet (2025): in $3, out $15, Claude Opus 4 (2025): in $15, out $75, Claude Sonnet 4 (2025): in $3, out $15, Gemini 2.0 Flash (2025): in $0.1, out $0.4, Gemini 2.0 Flash-Lite (2025): in $0.075, out $0.3, GPT-4.1 (2025): in $2, out $8, GPT-4.1 mini (2025): in $0.4, out $1.6, GPT-4.1 nano (2025): in $0.1, out $0.4, GPT-4.5 preview (2025): in $75, out $150, o3-mini (2025): in $1.1, out $4.4

Architecture mix

Dense transformer vs mixture-of-experts, across models with a disclosed architecture.

08
159tracked
  • Dense transformer143
  • Mixture of experts16

Dense transformer: 143, Mixture of experts: 16

Regional footprint

Share of tracked organisation and model activity by headquarters region.

09
  • North America1008
  • Asia453
  • Europe344
  • Middle East48

North America: 1008, Asia: 453, Europe: 344, Middle East: 48

Lab founding timeline

Founding year of tracked labs against how much tracked activity they have generated.

10

IBM: founded 1911, 47 records, Microsoft: founded 1975, 72 records, Apple: founded 1976, 20 records, NVIDIA: founded 1993, 24 records, Amazon: founded 1994, 73 records, Alibaba Cloud: founded 2009, 244 records, Snowflake: founded 2012, 13 records, Meta AI: founded 2013, 124 records, Databricks: founded 2013, 15 records, OpenAI: founded 2015, 253 records, Hugging Face: founded 2016, 13 records, AI21 Labs: founded 2017, 48 records, Cohere: founded 2019, 75 records, Zhipu AI: founded 2019, 43 records, Stability AI: founded 2019, 13 records, Anthropic: founded 2021, 176 records, MosaicML: founded 2021, 23 records, Inflection AI: founded 2022, 13 records, Reka AI: founded 2022, 25 records, Google DeepMind: founded 2023, 172 records, Mistral AI: founded 2023, 134 records, xAI: founded 2023, 67 records, DeepSeek: founded 2023, 123 records, 01.AI: founded 2023, 35 records

Model modality field

Declared input modalities. An empty cell means not documented, not unsupported.

11
ModelaudioimageNepali textneural recordingstextvideo
GPT-5.6 Sol
Muse Spark 1.1
Claude Sonnet 5
Gemini 3.5 Flash
TRIBE v2
NepaliGPT
Nepali GPT-2
Nepali RoBERTa
Nepali BERT

Matrix of model names and documented modalities.

Funding event scale

Disclosed historical funding rounds, square-root scaled for visibility.

12
  1. OpenAI
    $10,000M
  2. xAI
    $6,000M
  3. Anthropic
    $4,000M
  4. Inflection AI
    $1,300M
  5. Cohere
    $500M
  6. Mistral AI
    $415M
  7. Square-root scale preserves visibility across a wide funding range.

Cohere: $500M, xAI: $6000M, Mistral AI: $415M, Anthropic: $4000M, Inflection AI: $1300M, OpenAI: $10000M

Synthesis

The curated evidence shows three connected shifts: agentic multimodal models now arrive as tiered product families; AI governance is moving into implementation; and Nepal is building policy, education and Nepali-language research foundations while public evidence of scaled deployment and independent evaluation remains limited.

03 · Model direction

The release is no longer just a model checkpoint.

The latest tracked families package differentiated price-performance tiers, tool use, computer interaction and distribution APIs. That makes product governance and runtime observability as important as static benchmark tables.

Newest tracked model

GPT-5.6 Sol

2026-07-09 · reasoning and agentic

  1. GPT-5.6 Sol

    OpenAI · 2026-07-09 · frontier · context not disclosed in reviewed source

    Flagship member of the three-tier GPT-5.6 family; public context-window detail was not asserted in the reviewed launch source.

  2. Muse Spark 1.1

    Meta AI · 2026-07-09 · multimodal · context not disclosed in reviewed source

    Meta pairs the model with a first-party API preview, extending its distribution strategy beyond open-weight releases.

  3. Claude Sonnet 5

    Anthropic · 2026-06-30 · frontier · context not disclosed in reviewed source

    A scale-oriented Sonnet release announced alongside Claude Science.

  4. Gemini 3.5 Flash

    Google DeepMind · 2026-05-20 · multimodal · context not disclosed in reviewed source

    Generally available at Google I/O 2026 as the first Gemini 3.5 model.

  5. TRIBE v2

    Meta AI · 2026-03-26 · open-weight · context not disclosed in reviewed source

    A predictive foundation model for modelling human neural activity.

  6. NepaliGPT

    NepaliGPT research team · 2025-06-19 · open-weight · context not disclosed in reviewed source

    Important local research infrastructure; claims are deliberately bounded to what the paper reports.

  7. Nepali GPT-2

    Kathmandu University · 2024-11-24 · open-weight · context not disclosed in reviewed source

    One of three Nepali transformer families trained by KU researchers; the paper is the validated source of record.

  8. Nepali RoBERTa

    Kathmandu University · 2024-11-24 · open-weight · context not disclosed in reviewed source

    Encoder model from KU's 27.5 GB Nepali corpus programme.

  9. Nepali BERT

    Kathmandu University · 2024-11-24 · open-weight · context not disclosed in reviewed source

    A Nepali-specific encoder baseline intended for downstream language tasks.

Documented Nepal initiatives · newest evidence first

Department of Artificial Intelligence

Kathmandu University · Active

Provides a visible domestic talent and research pathway.

Purpose: Dedicated AI education and research

Availability: University programmes

Capabilities: BTech AI · MTech AI · MS and PhD research · industry collaboration

Limitations: Programme listings do not by themselves measure research impact or graduate outcomes

View source

National AI Policy 2082

Government of Nepal — MoCIT · Published 2025; implementation ongoing

Creates Nepal's first consolidated national AI direction.

Purpose: National AI governance, institutional coordination, skills and responsible adoption

Availability: Official policy PDF

Capabilities: AI Regulation Council framework · National AI Centre framework · AI Excellence Centre framework

Limitations: Publication is not implementation · Delivery milestones and measured outcomes require ongoing verification

View source

NepaliGPT

Independent NepaliGPT research team · 2025 research release

Creates dedicated generative and evaluation resources for a low-resource language.

Purpose: Nepali-language text generation and evaluation

Availability: Preprint available; model/deployment access not fully documented

Capabilities: Nepali text generation · Dedicated Devanagari corpus · 4,296-pair QA benchmark

Limitations: Preprint status · Limited independent evaluation · Unclear long-term hosting and weight availability

View source

Nepali BERT, RoBERTa and GPT-2 programme

Kathmandu University — ILPRL · 2024 preprint; 2025 peer-reviewed publication

Built a 27.5 GB Nepali corpus and stronger documented baselines.

Purpose: Nepali language understanding and generation baselines

Availability: Methods and results in peer-reviewed paper; artefact availability must be checked per model

Capabilities: Nepali classification and understanding · Text generation · Instruction-tuning research

Limitations: Legacy model architectures · Deployment documentation is incomplete · Compute and maintenance constraints

View source

05 · What to watch

Evidence gaps are part of the insight.

A trustworthy dashboard distinguishes what is known from what the ecosystem still needs to publish.

  1. 01No comparable public, continuously updated benchmark covering Nepali capability across global and local models.
  2. 02Public documentation of model licences, weights and reproducible training artefacts is inconsistent.
  3. 03Policy implementation outcomes and domestic compute capacity need measurable, regularly published indicators.

Continue from analysis to the underlying newest-first evidence.

Explore validated news