{
 "generated": "2026-10-06",
 "models": [
  {
   "id": "anthropic/claude-opus-5.5",
   "name": "Claude Opus 5.5",
   "provider": "Anthropic",
   "released": "2026-09-22",
   "in_price": 4.0,
   "out_price": 20.0,
   "ctx": 1000000,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (max)": {
     "value": 58,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "LMArena text Elo (opus-5.5-high, preliminary 4,552 votes)": {
     "value": 1504,
     "source": "https://arena.ai/leaderboard/text"
    },
    "Terminal-Bench 4.0": {
     "value": "66.4%",
     "source": "https://www.anthropic.com/news/claude-opus-5-5"
    },
    "FrontierCode v1.1": {
     "value": "54.4%",
     "source": "https://www.anthropic.com/news/claude-opus-5-5"
    },
    "CursorBench 4.0": {
     "value": "57.8%",
     "source": "https://www.anthropic.com/news/claude-opus-5-5"
    },
    "GDPval-AA v2.1": {
     "value": "1846 Elo",
     "source": "https://www.anthropic.com/news/claude-opus-5-5"
    },
    "OSWorld 2.1": {
     "value": "81.8%",
     "source": "https://www.anthropic.com/news/claude-opus-5-5"
    },
    "Humanity's Last Exam (with tools)": {
     "value": "67.7%",
     "source": "https://www.anthropic.com/news/claude-opus-5-5"
    },
    "SWE-bench Pro": {
     "value": "unverified",
     "source": "https://llm-stats.com/blog/research/claude-opus-5-5-launch"
    }
   },
   "notes": "List = OpenRouter $4/$20 (cache read $0.20, write $5). Context 1M from OpenRouter/AA (announcement does not state it). Anthropic: ~40% cheaper to run than Opus 5, >30% faster output. Secondary blogs claim SWE-bench Pro 89.9% (https://llm-stats.com/blog/research/claude-opus-5-5-launch) but the figure is not in the announcement HTML - not used.",
   "sources": [
    "https://www.anthropic.com/news/claude-opus-5-5",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models",
    "https://arena.ai/leaderboard/text"
   ]
  },
  {
   "id": "anthropic/claude-sonnet-5.5",
   "name": "Claude Sonnet 5.5",
   "provider": "Anthropic",
   "released": "2026-09-28",
   "in_price": 2.0,
   "out_price": 10.0,
   "ctx": 1000000,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (max)": {
     "value": 56,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "Terminal-Bench 4.0": {
     "value": "70.6%",
     "source": "https://www.anthropic.com/claude-sonnet-5-5"
    },
    "CursorBench 4.0": {
     "value": "55.5%",
     "source": "https://www.anthropic.com/claude-sonnet-5-5"
    },
    "FrontierCode 1.1 Main (max)": {
     "value": "46.2%",
     "source": "https://www.anthropic.com/claude-sonnet-5-5"
    },
    "GDPval-AA v2.1": {
     "value": "1844",
     "source": "https://www.anthropic.com/claude-sonnet-5-5"
    },
    "OSWorld 2.1": {
     "value": "80.1%",
     "source": "https://www.anthropic.com/claude-sonnet-5-5"
    },
    "Humanity's Last Exam": {
     "value": "64.5%",
     "source": "https://www.anthropic.com/claude-sonnet-5-5"
    },
    "LMArena text Elo (sonnet-5.5-xhigh, 3,145 votes, rank 45)": {
     "value": 1471,
     "source": "https://arena.ai/leaderboard/text"
    }
   },
   "notes": "Same price as Sonnet 5; Anthropic says 30%+ faster and up to 30% cheaper per task. Beats Opus 5.5 on Terminal-Bench 4.0 (70.6 vs 66.4).",
   "sources": [
    "https://www.anthropic.com/claude-sonnet-5-5",
    "https://www.anthropic.com/news",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  },
  {
   "id": "anthropic/claude-fable-5.1",
   "name": "Claude Fable 5.1",
   "provider": "Anthropic",
   "released": "2026-09-01",
   "in_price": 10.0,
   "out_price": 50.0,
   "ctx": 1000000,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (max)": {
     "value": 53,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "LMArena text Elo (fable-5.1-max)": {
     "value": 1501,
     "source": "https://arena.ai/leaderboard/text"
    },
    "Terminal-Bench 4.0": {
     "value": "55.8%",
     "source": "https://www.anthropic.com/claude-fable-and-mythos-5-1"
    },
    "Terminal-Bench-Science 0.1": {
     "value": "52.6%",
     "source": "https://www.anthropic.com/claude-fable-and-mythos-5-1"
    },
    "GDPval-AA v2": {
     "value": "1853",
     "source": "https://www.anthropic.com/claude-fable-and-mythos-5-1"
    },
    "OSWorld 2.0 (partial)": {
     "value": "77.9%",
     "source": "https://www.anthropic.com/claude-fable-and-mythos-5-1"
    },
    "Humanity's Last Exam (no tools)": {
     "value": "60.9%",
     "source": "https://www.anthropic.com/claude-fable-and-mythos-5-1"
    },
    "Humanity's Last Exam (with tools)": {
     "value": "65.0%",
     "source": "https://www.anthropic.com/claude-fable-and-mythos-5-1"
    },
    "CursorBench 3.2.0": {
     "value": "73.4%",
     "source": "https://www.anthropic.com/claude-fable-and-mythos-5-1"
    }
   },
   "notes": "Released alongside Mythos 5.1 (same model, trusted-access safeguards). Cache read cut to $0.25. Anthropic's own Opus 5.5 post lists Fable 5.1 HLE as 65.6% vs 65.0% on the Fable post - minor vendor inconsistency. Context 1M per OpenRouter.",
   "sources": [
    "https://www.anthropic.com/claude-fable-and-mythos-5-1",
    "https://www.anthropic.com/news",
    "https://openrouter.ai/api/v1/models"
   ]
  },
  {
   "id": "openai/gpt-6-astra",
   "name": "GPT-6 Astra",
   "provider": "OpenAI",
   "released": "2026-09-03",
   "in_price": 10.0,
   "out_price": 50.0,
   "ctx": 1050000,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (max)": {
     "value": 53,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "HealthBench Professional": {
     "value": "64.7",
     "source": "https://deploymentsafety.openai.com/gpt-6-astra"
    },
    "GPQA Diamond": {
     "value": "unverified",
     "source": "https://computingforgeeks.com/gpt-6-astra-released-features-benchmarks/"
    },
    "DeepSWE v1.1 (xhigh)": {
     "value": "unverified",
     "source": "https://computingforgeeks.com/gpt-6-sol-luna-released-features-benchmarks/"
    },
    "LMArena text Elo (astra-max, 9,156 votes, rank 29)": {
     "value": 1477,
     "source": "https://arena.ai/leaderboard/text"
    }
   },
   "notes": "System card published 2026-09-03; OpenRouter listed 2026-09-04. >272K-token prompts billed $20/$75 on OpenRouter. Max output 128K, knowledge cutoff 2026-04-30 (OpenAI models page). Secondary-only (openai.com returns 403): GPQA Diamond 96.0%, ARC-AGI-3 99.9% (adapter harness), DeepSWE v1.1 74.1% - https://computingforgeeks.com/gpt-6-astra-released-features-benchmarks/",
   "sources": [
    "https://deploymentsafety.openai.com/gpt-6-astra",
    "https://developers.openai.com/api/docs/models",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  },
  {
   "id": "openai/gpt-6-astra-pro",
   "name": "GPT-6 Astra Pro",
   "provider": "OpenAI",
   "released": "2026-09-04",
   "in_price": 10.0,
   "out_price": 50.0,
   "ctx": 1050000,
   "open_weights": false,
   "scores": {},
   "notes": "Same model as GPT-6 Astra served with reasoning.mode=pro; same per-token price but spends far more tokens (OpenRouter description). Not separately benchmarked; date = OpenRouter listing.",
   "sources": [
    "https://openrouter.ai/api/v1/models"
   ]
  },
  {
   "id": "openai/gpt-6-sol",
   "name": "GPT-6 Sol",
   "provider": "OpenAI",
   "released": "2026-09-22",
   "in_price": 2.0,
   "out_price": 10.0,
   "ctx": 1050000,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3": {
     "value": "unverified",
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "LMArena text Elo (sol-max, 8,787 votes, rank 64)": {
     "value": 1457,
     "source": "https://arena.ai/leaderboard/text"
    },
    "DeepSWE v1.1 (max)": {
     "value": "unverified",
     "source": "https://computingforgeeks.com/gpt-6-sol-luna-released-features-benchmarks/"
    }
   },
   "notes": "Released with Luna on 2026-09-22 (system-card appendix added that day). Superseded a week later by GPT-6.1 Sol at the same price. Cache read $0.20 vs $0.10 for 6.1 Sol on OpenRouter. AA v4.3 board lists GPT-6.1 Sol, not 6 Sol. Secondary-only: DeepSWE v1.1 68.8% (max).",
   "sources": [
    "https://deploymentsafety.openai.com/gpt-6-astra/sec:appendix-sol-luna",
    "https://openrouter.ai/api/v1/models"
   ]
  },
  {
   "id": "openai/gpt-6-sol-pro",
   "name": "GPT-6 Sol Pro",
   "provider": "OpenAI",
   "released": "2026-09-22",
   "in_price": 2.0,
   "out_price": 10.0,
   "ctx": 1050000,
   "open_weights": false,
   "scores": {},
   "notes": "GPT-6 Sol with reasoning.mode=pro (OpenRouter). Same per-token price, higher token use.",
   "sources": [
    "https://openrouter.ai/api/v1/models"
   ]
  },
  {
   "id": "openai/gpt-6-luna",
   "name": "GPT-6 Luna",
   "provider": "OpenAI",
   "released": "2026-09-22",
   "in_price": 0.1,
   "out_price": 0.5,
   "ctx": 1050000,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (max)": {
     "value": 38,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "LMArena text Elo (luna-max, rank 90)": {
     "value": 1443,
     "source": "https://arena.ai/leaderboard/text"
    }
   },
   "notes": "Cheapest GPT-6 tier. Knowledge cutoff 2026-05-18 (OpenAI models page). OpenRouter also lists gpt-6-luna-pro (same price, pro reasoning mode).",
   "sources": [
    "https://developers.openai.com/api/docs/models",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  },
  {
   "id": "openai/gpt-6.1-sol",
   "name": "GPT-6.1 Sol",
   "provider": "OpenAI",
   "released": "2026-09-29",
   "in_price": 2.0,
   "out_price": 10.0,
   "ctx": 1050000,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (max)": {
     "value": 52,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "LMArena text Elo (6.1-sol-max, 3,071 votes, rank 21)": {
     "value": 1483,
     "source": "https://arena.ai/leaderboard/text"
    },
    "DeepSWE v1.1 (high)": {
     "value": "unverified",
     "source": "https://computingforgeeks.com/gpt-6-sol-luna-released-features-benchmarks/"
    }
   },
   "notes": "OpenAI: 'nearly matches GPT-6 Astra' on agentic coding/computer use at one-fifth Astra's price. AA cost to run the index at max: $0.72 vs Astra $3.26. openai.com returned 403 to fetch; claims via search snippet of the official post. Secondary-only: DeepSWE v1.1 75.2% (high).",
   "sources": [
    "https://openai.com/index/introducing-gpt-6-1-sol/",
    "https://developers.openai.com/api/docs/models",
    "https://deploymentsafety.openai.com/gpt-6-1-sol",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  },
  {
   "id": "openai/gpt-6.1-sol-pro",
   "name": "GPT-6.1 Sol Pro",
   "provider": "OpenAI",
   "released": "2026-09-29",
   "in_price": 2.0,
   "out_price": 10.0,
   "ctx": 1050000,
   "open_weights": false,
   "scores": {},
   "notes": "GPT-6.1 Sol with reasoning.mode=pro (OpenRouter).",
   "sources": [
    "https://openrouter.ai/api/v1/models"
   ]
  },
  {
   "id": "google/gemini-3.8-flash",
   "name": "Gemini 3.8 Flash",
   "provider": "Google",
   "released": "2026-09-02",
   "in_price": 0.75,
   "out_price": 3.75,
   "ctx": 1048576,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (high)": {
     "value": 41,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "LMArena text Elo (3.8-flash-high, preliminary)": {
     "value": 1495,
     "source": "https://arena.ai/leaderboard/text"
    },
    "HLE-Verified": {
     "value": "54.9%",
     "source": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/"
    },
    "DeepSWE v1.1": {
     "value": "unverified",
     "source": "https://the-decoder.com/gemini-3-8-flash-is-googles-third-budget-model-in-six-weeks-while-frontier-models-remain-mia/"
    }
   },
   "notes": "Introductory price $0.75/$3.75 until 2026-12-31, then $1.50/$7.50 from 2027-01-01. Ships with a restricted 3.8 Flash Cyber variant (Fairwind Program). Third Flash release in six weeks. The Decoder reports DeepSWE v1.1 73.7% from Google's chart; Google's page text only says it 'outperforms most larger frontier models' (number likely in an image).",
   "sources": [
    "https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models",
    "https://arena.ai/leaderboard/text"
   ]
  },
  {
   "id": "google/gemini-4-argon",
   "name": "Gemini 4 Argon",
   "provider": "Google",
   "released": "2026-09-30",
   "in_price": 2.0,
   "out_price": 10.0,
   "ctx": 1000000,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (high)": {
     "value": 53,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "LMArena text Elo (argon-high, preliminary 4,932 votes) - #1": {
     "value": 1525,
     "source": "https://arena.ai/leaderboard/text"
    },
    "DeepSWE v1.1": {
     "value": "77.9%",
     "source": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"
    },
    "AutomationBench": {
     "value": "51.3%",
     "source": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"
    }
   },
   "notes": "Announced 2026-09-30; restricted rollout, not GA. This is Google's answer to 'did a Gemini 3.x Pro ship': no. No Gemini 3.x Pro since 3.1 Pro Preview (Feb 2026); Google skipped to Gemini 4 Argon. Only trusted cyber defenders (Fairwind Program) have it so far; not on OpenRouter; no public model ID. Introductory $2/$10, later $4/$20. 1M-token OUTPUT limit; input context assumed 1M per AA/Arena listing.",
   "sources": [
    "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/",
    "https://artificialanalysis.ai/leaderboards/models",
    "https://arena.ai/leaderboard/text",
    "https://the-decoder.com/gemini-3-8-flash-is-googles-third-budget-model-in-six-weeks-while-frontier-models-remain-mia/"
   ]
  },
  {
   "id": "x-ai/grok-4.7",
   "name": "Grok 4.7",
   "provider": "SpaceXAI (xAI)",
   "released": "2026-09-21",
   "in_price": 2.0,
   "out_price": 6.0,
   "ctx": 500000,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (xhigh)": {
     "value": 46,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "CursorBench 4.0": {
     "value": "46.3%",
     "source": "https://x.ai/news/grok-4-7"
    },
    "DeepSWE v1.1 (high)": {
     "value": "71.0%",
     "source": "https://x.ai/news/grok-4-7"
    },
    "Terminal-Bench 4.0": {
     "value": "37.6%",
     "source": "https://x.ai/news/grok-4-7"
    },
    "AA Briefcase v1.1": {
     "value": "1,657",
     "source": "https://x.ai/news/grok-4-7"
    },
    "LMArena text Elo (xhigh, 6,405 votes, rank 91)": {
     "value": 1442,
     "source": "https://arena.ai/leaderboard/text"
    }
   },
   "notes": "Same price as Grok 4.6; prompts >200K tokens billed $4/$12 on OpenRouter. Fast variant at 2x price. Secondary sites quote Terminal-Bench 38.0% - vendor page says 37.6%.",
   "sources": [
    "https://x.ai/news/grok-4-7",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  },
  {
   "id": "mistralai/mistral-large-4-0",
   "name": "Mistral Large 4 (Preview)",
   "provider": "Mistral AI",
   "released": "2026-10-06",
   "in_price": 0.68,
   "out_price": 2.09,
   "ctx": 524288,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (Preview)": {
     "value": 38,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "DeepSWE v1.1": {
     "value": "61.7%",
     "source": "https://mistral.ai/news/mistral-large-4"
    },
    "Terminal-Bench 4.0": {
     "value": "28.3%",
     "source": "https://mistral.ai/news/mistral-large-4"
    }
   },
   "notes": "Public preview 2026-10-06. Open weights PROMISED by end of October 2026 - not yet on Hugging Face; flip to true when published. Mistral list price $1.36/$4.18; OpenRouter currently $0.68/$2.09 (half list). 1T total / 49B active MoE, multimodal. Context 524,288 per OpenRouter/AA (announcement doesn't state).",
   "sources": [
    "https://mistral.ai/news/mistral-large-4",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  },
  {
   "id": "qwen/qwen3.8-max-prime",
   "name": "Qwen3.8 Max Prime",
   "provider": "Alibaba (Qwen)",
   "released": "2026-09-23",
   "in_price": 4.0,
   "out_price": 12.0,
   "ctx": 1000000,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (Qwen3.8 Max 0902, same model)": {
     "value": 45,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "LMArena text Elo (qwen3.8-max, same model)": {
     "value": 1482,
     "source": "https://arena.ai/leaderboard/text"
    }
   },
   "notes": "Higher-throughput SKU of Qwen3.8 Max (0902) at 2x the price ($2/$6). No Qwen blog/model card - documented only on OpenRouter. Treat scores as Qwen3.8 Max's.",
   "sources": [
    "https://openrouter.ai/api/v1/models",
    "https://datanorth.ai/news/qwen-releases-qwen3-8-max-prime",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  },
  {
   "id": "z-ai/glm-5.3-prime",
   "name": "GLM-5.3 Prime",
   "provider": "Z.ai",
   "released": "2026-09-23",
   "in_price": 2.8,
   "out_price": 8.8,
   "ctx": 1000000,
   "open_weights": true,
   "scores": {
    "AA Intelligence Index v4.3 (GLM-5.3 max, same model)": {
     "value": 45,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "LMArena text Elo (glm-5.3-max, same model)": {
     "value": 1478,
     "source": "https://arena.ai/leaderboard/text"
    }
   },
   "notes": "Prime is a hosted acceleration SKU; the underlying GLM-5.3 weights are open (zai-org/GLM-5.3 on HF since 2026-08-25, MIT per arena.ai). High-speed variant of GLM-5.3 (1.5-2x throughput). OpenRouter dated slug says 20260921, listing created 2026-09-23. No independent benchmarks for Prime itself.",
   "sources": [
    "https://openrouter.ai/api/v1/models",
    "https://huggingface.co/zai-org/GLM-5.3",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  },
  {
   "id": "deepseek/deepseek-v4.1-flash",
   "name": "DeepSeek V4.1 Flash",
   "provider": "DeepSeek",
   "released": "2026-09-10",
   "in_price": 0.021421,
   "out_price": 1.32,
   "ctx": 1048576,
   "open_weights": true,
   "scores": {
    "AA Intelligence Index v4.3 (max)": {
     "value": 39,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "GPQA Diamond": {
     "value": "90.9",
     "source": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash"
    },
    "DeepSWE v1.1": {
     "value": "74.2",
     "source": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash"
    },
    "Terminal-Bench 2.1": {
     "value": "90.6",
     "source": "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash"
    },
    "Humanity's Last Exam": {
     "value": "36.8",
     "source": "https://thenextweb.com/news/deepseek-v4-1-flash-launch-v4-pro-retired-price-cut"
    },
    "LMArena text Elo (max, rank 38)": {
     "value": 1474,
     "source": "https://arena.ai/leaderboard/text"
    }
   },
   "notes": "MIT license. 552B total, 8B active prefill / 16B active decode; new Causal Encoder-Decoder architecture. DeepSeek list: peak/off-peak pricing (off-peak output $0.60/M, peak double) - OpenRouter's $0.0214 input price is unexplained. Vendor numbers self-reported.",
   "sources": [
    "https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash",
    "https://thenextweb.com/news/deepseek-v4-1-flash-launch-v4-pro-retired-price-cut",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  },
  {
   "id": "~deepseek/deepseek-pro-latest",
   "name": "DeepSeek Pro Latest (alias)",
   "provider": "DeepSeek",
   "released": "2026-09-14",
   "in_price": 0.19,
   "out_price": 5.0,
   "ctx": 1048576,
   "open_weights": false,
   "scores": {},
   "notes": "Date = alias creation. open_weights n/a (alias). Not a new model: OpenRouter alias that 'always redirects to the latest DeepSeek Pro'. Since 2026-09-14 04:00 UTC DeepSeek routes all V4-Pro requests to V4.1-Flash until a V4.1-Pro ships. The $0.19/$5 OpenRouter price looks inconsistent - do not chart.",
   "sources": [
    "https://openrouter.ai/api/v1/models",
    "https://thenextweb.com/news/deepseek-v4-1-flash-launch-v4-pro-retired-price-cut"
   ]
  },
  {
   "id": "meta/muse-spark-1.3",
   "name": "Muse Spark 1.3",
   "provider": "Meta",
   "released": "2026-09-02",
   "in_price": 1.25,
   "out_price": 4.25,
   "ctx": 1048576,
   "open_weights": false,
   "scores": {
    "AA Intelligence Index v4.3 (max)": {
     "value": 48,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "AA Intelligence Index v4.3 (xhigh)": {
     "value": 45,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "LMArena text Elo (max)": {
     "value": 1494,
     "source": "https://arena.ai/leaderboard/text"
    },
    "DeepSWE": {
     "value": "75.4% (Meta figure, max mode, via secondary)",
     "source": "https://www.eesel.ai/blog/muse-spark-1-3"
    }
   },
   "notes": "Closed weights. AA's launch article (2026-09-02) scored it 61 (xhigh) / 62 (max) on the PREVIOUS index version - not comparable to the v4.3 numbers. Max mode was limited preview at launch. A $0.10/$0.20 'contributor' variant also exists.",
   "sources": [
    "https://artificialanalysis.ai/articles/muse-spark-1-3",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models",
    "https://arena.ai/leaderboard/text"
   ]
  },
  {
   "id": "xiaomi/mimo-v2.6-pro",
   "name": "MiMo-V2.6-Pro",
   "provider": "Xiaomi",
   "released": "2026-09-22",
   "in_price": 0.435,
   "out_price": 0.87,
   "ctx": 1050000,
   "open_weights": true,
   "scores": {
    "AA Intelligence Index v4.3": {
     "value": 46,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    },
    "DeepSWE v1.1": {
     "value": "71.9",
     "source": "https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL"
    },
    "Terminal-Bench 4.0": {
     "value": "34.9",
     "source": "https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL"
    },
    "OSWorld-Verified": {
     "value": "82.0",
     "source": "https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL"
    },
    "LMArena text Elo (rank 26)": {
     "value": 1480,
     "source": "https://arena.ai/leaderboard/text"
    }
   },
   "notes": "MIT license, 1.02T total / 42B active, omnimodal. Highest-ranked open-weights model on AA v4.3 (ties Grok 4.7). OpenRouter listed 2026-09-21; Xiaomi release 2026-09-22.",
   "sources": [
    "https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL",
    "https://technode.com/2026/09/22/xiaomi-open-sources-mimo-v2-6-models-after-scaling-reinforcement-learning/",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  },
  {
   "id": "cohere/command-a-plus",
   "name": "Command A+",
   "provider": "Cohere",
   "released": "2026-05-20",
   "in_price": 0.3,
   "out_price": 1.5,
   "ctx": 192000,
   "open_weights": "unverified",
   "scores": {
    "AA Intelligence Index v4.3": {
     "value": 13,
     "source": "https://artificialanalysis.ai/leaderboards/models"
    }
   },
   "notes": "Cohere changelog says Apache 2.0 but no public HF repo confirmed. Announced 2026-05-20; OpenRouter listing 2026-09-22. Outside the Aug 15-Oct 6 window as a model release - only newly listed on OpenRouter. 218B total / 25B active MoE. Cohere: 128K input + 64K output (=192K total, matching OpenRouter). Low AA score likely reflects a non-reasoning run - verify before featuring.",
   "sources": [
    "https://docs.cohere.com/changelog/command-a-plus-05-2026",
    "https://openrouter.ai/api/v1/models",
    "https://artificialanalysis.ai/leaderboards/models"
   ]
  }
 ],
 "arena_top": [
  {
   "rank": 1,
   "model": "gemini-4-argon-high (preliminary)",
   "elo": 1525,
   "source": "https://arena.ai/leaderboard/text (board updated 2026-10-02, accessed 2026-10-06)"
  },
  {
   "rank": 2,
   "model": "claude-opus-4-6-high",
   "elo": 1505,
   "source": "https://arena.ai/leaderboard/text (board updated 2026-10-02, accessed 2026-10-06)"
  },
  {
   "rank": 3,
   "model": "claude-fable-5-high",
   "elo": 1504,
   "source": "https://arena.ai/leaderboard/text (board updated 2026-10-02, accessed 2026-10-06)"
  },
  {
   "rank": 4,
   "model": "claude-opus-5.5-high",
   "elo": 1504,
   "source": "https://arena.ai/leaderboard/text (board updated 2026-10-02, accessed 2026-10-06)"
  },
  {
   "rank": 5,
   "model": "claude-opus-4-7-high",
   "elo": 1501,
   "source": "https://arena.ai/leaderboard/text (board updated 2026-10-02, accessed 2026-10-06)"
  },
  {
   "rank": 6,
   "model": "claude-fable-5.1-max",
   "elo": 1501,
   "source": "https://arena.ai/leaderboard/text (board updated 2026-10-02, accessed 2026-10-06)"
  },
  {
   "rank": 7,
   "model": "claude-opus-4-6",
   "elo": 1497,
   "source": "https://arena.ai/leaderboard/text (board updated 2026-10-02, accessed 2026-10-06)"
  },
  {
   "rank": 8,
   "model": "gemini-3.8-flash-high (preliminary)",
   "elo": 1495,
   "source": "https://arena.ai/leaderboard/text (board updated 2026-10-02, accessed 2026-10-06)"
  },
  {
   "rank": 9,
   "model": "muse-spark-1.3-max",
   "elo": 1494,
   "source": "https://arena.ai/leaderboard/text (board updated 2026-10-02, accessed 2026-10-06)"
  },
  {
   "rank": 10,
   "model": "claude-opus-4-7",
   "elo": 1494,
   "source": "https://arena.ai/leaderboard/text (board updated 2026-10-02, accessed 2026-10-06)"
  }
 ],
 "aa_top": [
  {
   "model": "Claude Opus 5.5 (max)",
   "index": 58,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  },
  {
   "model": "Claude Sonnet 5.5 (max)",
   "index": 56,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  },
  {
   "model": "Claude Fable 5.1 (max)",
   "index": 53,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  },
  {
   "model": "GPT-6 Astra (max)",
   "index": 53,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  },
  {
   "model": "Gemini 4 Argon (high)",
   "index": 53,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  },
  {
   "model": "GPT-6.1 Sol (max)",
   "index": 52,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  },
  {
   "model": "Muse Spark 1.3 (max)",
   "index": 48,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  },
  {
   "model": "Grok 4.7 (xhigh)",
   "index": 46,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  },
  {
   "model": "MiMo-V2.6-Pro",
   "index": 46,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  },
  {
   "model": "Qwen3.8 Max (0902)",
   "index": 45,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  },
  {
   "model": "GLM-5.3 (max)",
   "index": 45,
   "source": "https://artificialanalysis.ai/leaderboards/models (Intelligence Index v4.3, best setting per model, accessed 2026-10-06)"
  }
 ],
 "routers": [
  {
   "id": "nvidia/switchyard",
   "what_it_is": "NVIDIA NeMo Switchyard: open-source (Rust) model router that sends each agent step to a cheap or a frontier model to cut cost. On OpenRouter it defaults to OpenRouter market data to pick the most popular low- and high-cost models; you can pass two models and a routing algorithm. NVIDIA claims mixing open models with Opus 4.8 kept frontier accuracy at ~1/3 the cost (vendor claim). OpenRouter listing 2026-09-21; price shown as variable (-1) on OpenRouter, so presumably billed at the routed model's rate (inferred, not documented).",
   "underlying_models": "user-defined pool; default = popular low-cost + high-cost models from OpenRouter data (e.g., Nemotron 3.5 Lightning locally)",
   "source": "https://kdnuggets.com/switchyard-nvidias-open-source-routing-library ; https://openrouter.ai/api/v1/models"
  },
  {
   "id": "typesafe/jev-router",
   "what_it_is": "Router from TypeSafe that picks the model AND reasoning effort per request, built on Jev. Jev is TypeSafe's first 'System One' model: a decision model, not an LLM - it returns typed choices/scores with a calibrated probability in one pass (~151 ms median), billed on input only (~$0.042/M input per OrcaRouter). Jev 1.13 launched 2026-09-15; Jev Router listed on OpenRouter 2026-09-25.",
   "underlying_models": "not disclosed (OpenRouter catalog models; picks model + effort)",
   "source": "https://openrouter.ai/docs/guides/community/typesafe-sdk ; https://www.orcarouter.ai/blog/jev-system-one ; https://openrouter.ai/api/v1/models"
  },
  {
   "id": "unbiased/pareto",
   "what_it_is": "Described by OpenRouter only as a 'multimodal composite model' for research, coding and agentic work. No public info on the company 'Unbiased' or which models it composes. pareto listed 2026-09-17 ($2.50/$7.50, 262K ctx); pareto-26.10-preview listed 2026-10-01 ($0.80/$3.20, 1M ctx). Not to be confused with openrouter/pareto-code (OpenRouter's own coding router ranked by AA coding percentiles).",
   "underlying_models": "unverified / undisclosed",
   "source": "https://openrouter.ai/unbiased ; https://openrouter.ai/api/v1/models"
  },
  {
   "id": "openrouter/auto",
   "what_it_is": "OpenRouter's own Auto Router: classifies the prompt into ~30 task types, ranks models by what the community spends on that task type over a trailing 7 days, filters by cost_tier (low...max). No extra fee - you pay the selected model's rate. 2M ctx ceiling. openrouter/auto-beta (2026-07-17) is the experimental version.",
   "underlying_models": "full OpenRouter catalog (restrict with allowed_models / excluded_models)",
   "source": "https://openrouter.ai/docs/guides/routing/routers/auto-router"
  }
 ],
 "caveats": [
  "AA Intelligence Index is now v4.3 (adds AutomationBench-AA, drops tau3-Banking, Terminal-Bench moves to 4.0). Scores are much lower than older versions (e.g., Muse Spark 1.3 was 61-62 at launch, Gemini 3.8 Flash 59 per The Decoder; now 48 and 41). Do NOT mix with numbers on the Aug 29 page without re-pulling all rows.",
  "Arena scores marked 'preliminary' (Gemini 4 Argon, Gemini 3.8 Flash) have few votes and wide CIs (+/-9 for Argon and Opus 5.5). Opus 5.5 at 1504 is statistically tied with ranks 2-6. Arena ranks of the new OpenAI models are far lower than their AA ranks (6.1 Sol #21, Astra #29, Sol #64, Luna #90); Sonnet 5.5 #45, Grok 4.7 #91; Mistral Large 4 and Command A+ not on the board.",
  "openai.com returned HTTP 403 to automated fetch; OpenAI benchmark numbers (GPQA, DeepSWE, ARC-AGI-3) seen only in secondary coverage and are marked unverified. Prices/context come from developers.openai.com and OpenRouter.",
  "Most vendor benchmarks have moved to new suites (DeepSWE v1.1, Terminal-Bench 4.0, OSWorld 2.x, CursorBench 4.0). SWE-bench Verified and ARC-AGI numbers were not found in any primary source for these models; GPQA only for DeepSeek V4.1 Flash (vendor).",
  "No Gemini 3.x Pro shipped in the window. Google's new flagship is Gemini 4 Argon (2026-09-30), restricted to the Fairwind Program; not on OpenRouter.",
  "OpenRouter 'released' dates are listing-creation dates and can differ by a day from vendor announcements (Astra 09-03 vs 09-04, MiMo 09-22 vs 09-21).",
  "Pro variants of GPT-6 (Astra/Sol/Luna/6.1 Sol Pro) are the same weights in reasoning.mode=pro at the same per-token price; cost per task is much higher.",
  "Prices are base tier; OpenAI and xAI charge higher rates above 272K / 200K prompt tokens on OpenRouter. Google's Gemini 3.8 Flash and 4 Argon prices are introductory and double later.",
  "Model names like SpaceXAI (formerly xAI) are as listed by OpenRouter/AA."
 ]
}