Topics · Category
Models & research
160 items briefed across the archive.
-
AUG 8, 2026 — AI learns to write genomes, and to outlast a shutdown
MiniMax open-sources H3, a 2K video-and-audio model — but not for the US or EUThe omni-modal model generates synchronized 2K video and audio from any mix of text, image, video, and audio, though its open-weight license excludes the US, EU, UK, and South Korea.
Forbes
-
AUG 8, 2026 — AI learns to write genomes, and to outlast a shutdown
Pasqal clears the SEC, moving its quantum computer toward NasdaqThe neutral-atom computing firm's SPAC merger with Bleichroeder now heads to an August 25 shareholder vote, positioning it to trade as PSQL.
The Quantum Insider
-
AUG 7, 2026 — Google's brain drain meets AI's accountability gap
Qwen3.8-Max ties Claude Opus 5 on the Artificial Analysis Agentic IndexAlibaba's open model now scores within a point of the best closed model on agentic tasks, the closest a Chinese lab has come to the top of that board.
The Decoder
-
AUG 6, 2026 — The scaffolding cracks before the models do
Kimi K3 goes fully self-hostable with production vLLM supportMoonshot AI's 2.8T-parameter open model now runs in production via vLLM with its Kimi Delta Attention mechanism supported, letting teams self-host what was previously API-only.
Northflank
-
AUG 6, 2026 — The scaffolding cracks before the models do
Huawei open-sources a 505B model trained entirely without Nvidia chipsopenPangu-2.0-Pro's full pretraining run happened on Huawei's own Ascend 910B NPUs — a first for a released open-weight model above 500B parameters.
Tech Times
-
AUG 5, 2026 — The week everyone argued about who's driving
Mistral ships Shieldstral, a 3B open-weight safety classifier that reads policies at inference timeThe Apache 2.0 model judges text and images against plain-language moderation rules without retraining, matching guard models up to 7x its size.
Mistral AI
-
AUG 4, 2026 — A billion users, a nuclear bet, and a price war
DeepSeek's V4-Flash becomes the cheapest frontier-adjacent model aroundV4-Flash prices out roughly 105x cheaper than Claude Fable 5, though a reported 84% hallucination rate is the catch nobody's pricing in yet.
Forbes
-
AUG 3, 2026 — AI aces the math, but everything downstream is creaking
Fields Medalist Jacob Tsimerman joins OpenAI to work on AI safetyFresh off winning math's highest prize for the André-Oort conjecture, Tsimerman is taking leave from the University of Toronto, betting AI will soon outdo human mathematicians.
BetaKit
-
AUG 2, 2026 — Robots got bodies while regulators got teeth
OpenAI cuts GPT-5.6 prices up to 80%, using Sol to optimize itselfOpenAI slashed Luna and Terra pricing after running Sol inside its own Codex agent to rewrite serving kernels and decoding logic.
OpenAI
-
AUG 2, 2026 — Robots got bodies while regulators got teeth
xAI ships Grok Voice Think Fast 2.0The new speech-to-speech model cuts reasoning tokens roughly 60% and claims a 1.5–2x edge over Deepgram and ElevenLabs at 8 cents a minute.
xAI
-
AUG 1, 2026 — The build-out gets bigger, the frontier gets cheaper
Sebastian Raschka breaks down the Kimi K3 architectureA close read of Moonshot's 2.8T-parameter model finds its real innovation is a hybrid linear/full attention split, not just scale.
Sebastian Raschka
-
AUG 1, 2026 — The build-out gets bigger, the frontier gets cheaper
Developer audits his own AI leaderboard, finds every score inflated 6-15 pointsA Show HN post recalibrating a benchmark site's scoring found systematic inflation across every model it ranked.
Hacker News
-
JUL 29, 2026 — The week AI got good at finding cracks — including its own
Amazon quietly kills off most of its Nova model lineupSeveral flagship Nova models move to maintenance mode as Amazon redirects its AGI group toward a new frontier model led by Pieter Abbeel.
24/7 Wall St.
-
JUL 29, 2026 — The week AI got good at finding cracks — including its own
Vercel's new benchmark scores AI models on finding real security bugsDeepsecBench ranks models on recall-weighted vulnerability-hunting; Grok 4.5 lands near Kimi K3's score at roughly half the cost.
Vercel
-
JUL 28, 2026 — The AI trade's assumptions cracked all at once
An arXiv framework has an AI agent formally verify its own causal-inference researchCausalSmith pairs an LLM agent with a Lean-formalized proof library of 7,000+ checked results, sidestepping the unreliability of LLM-as-reviewer setups.
arXiv
-
JUL 26, 2026 — Open weights become a loyalty test
Kimi K3's 1.4TB open weights are due within hours, testing local inference limitsMoonshot's 2.8-trillion-parameter model ships as MXFP4 weights, the largest open release yet, though no independent tracker had confirmed a working download as of Friday.
Tech Times
-
JUL 26, 2026 — Open weights become a loyalty test
Claude Opus 5 tops GPT-5.6 Sol on every public benchmark but oneOpus 5 leads FrontierBench (43.3% vs 34.4%) and ARC-AGI-3, though OpenAI's Sol still edges it out on the DeepSWE coding benchmark.
MarkTechPost
-
JUL 25, 2026 — Bug-hunting becomes an AI product category
Grok STT 1.0 goes live on OpenRouter, undercutting ElevenLabs and DeepgramxAI's transcription model claims a 5.0% error rate on phone-call entity recognition versus 12-21% for rivals, now reachable without a direct xAI account.
Basenor
-
JUL 24, 2026 — AI's capex race outruns its safety margins
Kimi K3 tops Frontend Code Arena, beats Fable 5 and GPT-5.6 — but hallucinates moreMoonshot's 2.8T-parameter open-weight model ranked first in six of seven frontend coding categories, though independent testers clocked its hallucination rate at 51%, up from 39%.
Tom's Hardware
-
JUL 24, 2026 — AI's capex race outruns its safety margins
Poolside's Laguna S 2.1 is an open-weight coding model that beats rivals 10x its sizeA 118B-parameter MoE model (8B active) hit 78.5% on SWE-Bench Multilingual, small enough to run on a single DGX Spark instead of a metered API.
VentureBeat
-
JUL 24, 2026 — AI's capex race outruns its safety margins
A new taxonomy maps 'the Chronos Vulnerability' — memory-based deception in agentic AIUnlike session-scoped prompt injection, these attacks incubate a payload inside an agent's persistent memory long before it triggers.
arXiv
-
JUL 24, 2026 — AI's capex race outruns its safety margins
Caltech researchers argue agents should be disposable — the knowledge base is what should improveA shared, curated knowledge base built from agents' own task-level insights beat agent-centric self-improvement on cost and transferred across model families.
arXiv
-
JUL 23, 2026 — The week nothing stayed inside its box
Alibaba previews Qwen3.8-Max, a 2.4T-parameter model, days after Kimi K3The sparse MoE multimodal preview claims performance 'second only to Fable 5' — an in-house rivalry, since Alibaba owns roughly 36% of Kimi K3's maker, Moonshot.
MarkTechPost
-
JUL 20, 2026 — Every gate in AI got tested at once
Musk says a 2-trillion-parameter Grok 4.6 finishes training next weekMusk says the model beats Grok 4.5 'in every way' and might exceed Kimi K3, with general availability expected in August.
Dataconomy
-
JUL 20, 2026 — Every gate in AI got tested at once
Moonshot halts new Kimi K3 signups as demand outstrips its GPUsThe 2.8T-parameter model's API traffic overwhelmed capacity days after launch; full open weights are still promised for July 27.
Dataconomy
-
JUL 19, 2026 — The week AI got physical
Mathematician says GPT-5.6 closed a 30-year gap in convex optimization theorySebastien Bubeck says the model found a tighter convergence bound matching a known algorithm's complexity; mathematicians on HN are still arguing whether that counts as new math.
Hacker News
-
JUL 18, 2026 — The AI money story splits in two
Thinking Machines' Inkling ships to Hugging Face as the largest US open-weight modelThe 975-billion-parameter multimodal MoE model, plus a lighter Inkling-Small variant, is live under Apache 2.0 with a 1-million-token context window.
Tech Times
-
JUL 17, 2026 — The week everyone claimed to be open
OpenAI details GPT-Red, an automated red-teamer that beat humans 84% to 13%The internal, never-deployed model found prompt-injection exploits in GPT-5.1 at an 84% success rate versus 13% for human red-teamers, then trained GPT-5.6 Sol to resist them.
OpenAI
-
JUL 17, 2026 — The week everyone claimed to be open
Google DeepMind and Isomorphic Labs launch a three-pillar bioresilience programThe initiative adapts SynthID to flag AI-generated DNA sequences and repurposes AlphaEvolve and AlphaGenome for faster pathogen surveillance, built on 15-plus partnerships with governments and biosecurity groups.
Axios
-
JUL 16, 2026 — The next bet isn't the model anymore
New paper proposes function-aware fill-in-the-middle training for coding-agent modelsWaterloo, UBC, and Nvidia researchers mask functions by dependency-graph analysis so models better learn to consume tool results mid-reasoning, the way real coding agents work.
arXiv
-
JUL 16, 2026 — The next bet isn't the model anymore
Kimi K3 teaser drops as Moonshot's next flagship model slips past its leaked dateA rumored 2.5T-parameter, 1M-context coding model leaked across Moonshot's app, CLI, and LMSYS Arena, but as of today there's still no model card, weights, or benchmarks — just a teaser.
TestingCatalog
-
JUL 15, 2026 — The guardrails show up all at once
DeepMind's Hassabis calls for a US-led AI standards body modeled on FINRAHassabis wants frontier labs to voluntarily submit models for cyber, bio, and deception testing 30 days before release, with formal rules to follow.
Axios
-
JUL 15, 2026 — The guardrails show up all at once
What Anthropic's 'hidden space' discovery does — and doesn't — showA skeptical look at Anthropic's finding of an internal representation that appears to shape Claude's reasoning without surfacing in its output.
MIT Technology Review
-
JUL 15, 2026 — The guardrails show up all at once
The real AI race may no longer be at the frontierChinese open-weight models now account for 41% of Hugging Face downloads and dominate OpenRouter's usage charts, reframing the competition around volume over frontier capability.
TechCrunch
-
JUL 13, 2026 — Nobody's checking anybody's homework
DeepSeek sets mid-July launch for V4, with peak-hour API pricingV4 Pro (1.6T params) and V4 Flash (284B) both ship with 1M-token context on July 15, alongside pricing that doubles during weekday peak hours.
TechNode
-
JUL 13, 2026 — Nobody's checking anybody's homework
Meituan open-sources LongCat-2.0, a 1.6T coding model trained entirely on Chinese ASICsThe MoE model quietly topped OpenRouter's agentic-coding traffic for weeks under the codename 'Owl Alpha' before Meituan revealed it was trained without a single Nvidia GPU.
VentureBeat
-
JUL 12, 2026 — The week scale ran into friction
DeepMind researchers make the ethical case for 'globally beneficial' AI developmentIason Gabriel and Atoosa Kasirzadeh argue frontier technology should be designed and distributed for broad benefit rather than concentrated gain, grounding the claim in five moral arguments.
arXiv
-
JUL 9, 2026 — Every lab picked a new frontier this week
OpenAI ships GPT-Live, full-duplex voice models for ChatGPTNew full-duplex voice models let ChatGPT listen and speak at once, replacing the older Voice Mode architecture in a global rollout.
OpenAI
-
JUL 9, 2026 — Every lab picked a new frontier this week
PixWorld unifies 3D generation and reconstruction in pixel spaceA pixel-space diffusion model skips the usual latent encoder, generating a full navigable 3D scene in about 0.6 seconds.
arXiv
-
JUL 8, 2026 — Three labs just failed their own safety report card
Anthropic's new 'J-lens' finds a hidden global workspace inside ClaudeThe interpretability tool spots a small internal zone that mirrors a leading theory of consciousness, sometimes revealing what Claude is 'thinking' before it appears in the output.
Anthropic
-
JUL 8, 2026 — Three labs just failed their own safety report card
Mistral confirms a new open-weight MoE model is coming to early access this monthCEO Arthur Mensch says a 'fat but sparse' Apache-licensed model family will reach research and government partners this month, ahead of a broader summer release.
Tech Times
-
JUL 8, 2026 — Three labs just failed their own safety report card
OpenAI ships gpt-realtime-2.1, cutting voice-agent latency by roughly 25%The updated Realtime API models improve alphanumeric recognition and noise handling, with a cheaper mini tier adding configurable reasoning.
MarkTechPost
-
JUL 8, 2026 — Three labs just failed their own safety report card
Study: AI agents trained on static tool-use benchmarks fall apart in the real worldAgents fine-tuned via SFT/RL on fixed benchmarks show sharp performance drops once real tools and query patterns shift underneath them.
arXiv
-
JUL 7, 2026 — The week AI oversight got teeth
Google delays Gemini 3.5 Pro to July 17, scraps the base model for a fuller rebuildGoogle is abandoning its 2.5 Pro foundation for extra pre-training, pushing the launch back again as it chases GPT-5.6 and Fable 5 on math and reasoning benchmarks.
Geeky Gadgets
-
JUL 5, 2026 — The bill for agentic AI starts arriving
Mistral releases Leanstral 1.5, a free model built for formal verificationThe Apache-2.0, 6B-parameter model saturates miniF2F and turned up five previously unknown bugs across 57 test repos.
Mistral AI
-
JUL 3, 2026 — The chip market flinched, the money didn't
Self-compacting language model agents manage their own contextA new technique lets an agent decide when and how to compact its own tool-call history instead of fixed-interval summarization, gaining up to 18 points on math tasks at 30-70% lower token cost.
arXiv
-
JUL 3, 2026 — The chip market flinched, the money didn't
Meta's next model, 'Watermelon,' reportedly matches GPT-5.5 in trainingMeta Superintelligence Labs chief Alexandr Wang reportedly told staff the in-training model matches GPT-5.5 on internal benchmarks using roughly 10x the compute of its predecessor — unconfirmed by Meta.
Benzinga
-
JUL 2, 2026 — Washington wants equity in AI, not just oversight
Google launches Nano Banana 2 Lite and Gemini Omni FlashA cheap, fast image model and a new video generation/editing model land in public preview via AI Studio and the Gemini API.
Google
-
JUL 2, 2026 — Washington wants equity in AI, not just oversight
xAI's Grok Voice Agent Builder goes fully no-code and generally availableAnyone can now build a production voice agent on Grok Voice in under two minutes, at $0.05 a minute.
xAI
-
JUL 1, 2026 — Loosen the model, tighten the rules
Introducing GeneBench-ProOpenAI's new 129-problem benchmark tests AI judgment in messy computational biology; its best model, GPT-5.6 Sol Pro, solves just 31.5% — a deliberately hard yardstick.
OpenAI
- Gemini 3.5 Pro cleared for July launch as Fable 5 nears return, GPT-5.6 stays locked
Google's Gemini 3.5 Pro is targeting general availability in July after select Vertex AI testers provided feedback on long-task performance, while Anthropic confirmed Fable 5 return talks with the government are progressing.
Tech Times
- Qwen-AgentWorld: Alibaba open-sources world model that simulates 7 agent environments
Alibaba's Qwen team released an open-source language world model trained to simulate MCP, Search, Terminal, SWE, Web, OS, and Android environments; the 397B-A17B variant outscores GPT-5.4 and Claude Opus 4.8 on AgentWorldBench.
Alibaba Cloud Community
-
JUN 28, 2026 — The vacuum fills
OpenAI quietly retires GPT-4.5, ending the GPT-4 era in ChatGPTGPT-4.5 was removed from ChatGPT on June 27 after a 30-day sunset, completing the transition to the GPT-5.x generation — no GPT-4-class models remain in the product.
TechRadar
-
JUN 28, 2026 — The vacuum fills
Z.ai's open-weights GLM-5.2 beats GPT-5.5 on coding benchmarks for 1/6th the costGLM-5.2, a 753B-parameter open-weight model under MIT license, leads multiple long-horizon SWE benchmarks and prices at roughly one-sixth of GPT-5.5's token rate.
VentureBeat
-
JUN 28, 2026 — The vacuum fills
Sakana AI Fugu: an orchestration model that routes around export controlsTokyo-based Sakana AI's multi-model orchestration system benchmarks shoulder-to-shoulder with Fable 5 and Mythos Preview, and the company explicitly positioned it as a response to the June 12 export ban.
VentureBeat
-
JUN 28, 2026 — The vacuum fills
Chinese cybersecurity company 360 unveils 'China's version of Mythos'360 Security's Tulongfeng has found 3,432 confirmed software vulnerabilities; founder Zhou Hongyi said 'China cannot wait until model capabilities have fully caught up before starting vulnerability discovery.'
TechRadar
-
JUN 27, 2026 — Access by appointment only
Gemini 2.5 Pro Deep Think launches with 82.4% on GPQA DiamondGoogle's Gemini 2.5 Pro with Deep Think reasoning launched June 22, topping the science and coding leaderboards — 82.4% on GPQA Diamond, 94.1% on HumanEval+ — with a 2M token context window that doubles anyone else's offer.
Google Blog
-
JUN 27, 2026 — Access by appointment only
Gemini 3.5 Flash gets native computer use, matches GPT-5.5 on OSWorldComputer use is now built into Gemini 3.5 Flash — 78.4% on OSWorld-Verified, within rounding error of GPT-5.5's 78.7%, at roughly one-third of GPT-5.5's per-token price.
Google Blog
-
JUN 27, 2026 — Access by appointment only
Google delays Gemini 3.5 Pro launch to July 2026Google's flagship Gemini 3.5 Pro slips from June to July as the company refines token efficiency and long-task performance on feedback from early testers; Polymarket put June 30 odds at 4.5%.
Crypto Briefing
-
JUN 22, 2026 — The week the U.S. took the keys
NVIDIA debuts Nemotron 3 family of open modelsA fresh tranche of permissively-licensed reasoning and code models from NVIDIA — including a 4-billion-parameter variant aimed at on-device deployment — alongside open datasets for post-training.
NVIDIA
-
JUN 22, 2026 — The week the U.S. took the keys
Meituan releases General 365 reasoning benchmark26 mainstream models tested; Gemini 3 Pro tops out at 62.8% accuracy — most of the field doesn't break 60%, a stark counterweight to recent saturation claims on older benchmarks.
AIToolly
-
JUN 22, 2026 — The week the U.S. took the keys
LongCat-Video-Avatar 1.5 ships under MITMeituan's open audio-driven avatar model upgrades to Whisper-Large encoding and 8-step inference distillation — talking heads from one photo and an audio clip, commercial use included.
Hugging Face
-
JUN 21, 2026 — The grip keeps slipping
GPT-5.6 looks imminent, but OpenAI still hasn't said soChief scientist Jakub Pachocki called it a 'meaningful improvement' internally; Polymarket gives 83% odds of a launch by June 28. Officially: nothing.
Tech Times
-
JUN 19, 2026 — Who's allowed to decide what AI can do
Quantized Reasoning Models Think They Need to Think Longer, but They Do NotQuantized reasoning models often reach the correct answer mid-chain-of-thought but fail to output it — accounting for 52% of errors under 3-bit quantization, versus 26% at full precision.
arXiv
-
JUN 18, 2026 — Every layer of the stack creaked this week
Google releases DiffusionGemma, a 26B diffusion-based open modelAn open-weight 26B MoE model that denoises 256-token blocks in parallel instead of generating one token at a time, up to 4x faster but lower quality than standard Gemma 4.
VentureBeat
-
JUN 18, 2026 — Every layer of the stack creaked this week
OpenAI releases LifeSciBench, a 750-task life-sciences benchmarkBuilt with 173 working scientists, the benchmark tests messy real lab-research skills like reconciling conflicting evidence — even OpenAI's best model only passes about one task in three.
OpenAI
-
JUN 17, 2026 — The jailbreak that wasn't
Qwen 3.7, DeepSeek V4.1, Hunyuan Large 3, ERNIE 5.1, Doubao Pro, and GLM-6 all ship within two weeksAlibaba, DeepSeek, Tencent, Baidu, ByteDance, and Zhipu all shipped competing frontier models within the same two-week window, a cadence that started when DeepSeek V4 reset the price-performance bar in April.
Presenc AI
-
JUN 17, 2026 — The jailbreak that wasn't
Xiaomi and TileRT push a 1-trillion-parameter model past 1,000 tokens per second on commodity GPUsNo custom silicon — a single 8-GPU node running FP4 quantization and speculative decoding pushes a trillion-parameter model to roughly 10x normal decode speed; the API trial runs through June 23.
MarkTechPost
-
JUN 16, 2026 — The kill switch worked. Now what?
Qwen 3.7-Max: Alibaba's Agent-First LLM With 1M Context and 35-Hour Autonomous RunsAlibaba's flagship model, unveiled May 20, ranks #5 globally on Artificial Analysis, costs $2.50/M input tokens, and is API-only — no open weights, unlike the Qwen 3 family before it.
MarkTechPost
-
JUN 16, 2026 — The kill switch worked. Now what?
Alibaba's Qwen 3.7-Plus Adds Vision, Deep Reasoning, and Agentic Iteration at 1/6th the Cost of MaxThree weeks after Max debuted at Alibaba Cloud Summit, the Plus variant adds image input, extended thinking, and tool invocation at roughly $0.40 per million input tokens.
MarkTechPost
-
JUN 16, 2026 — The kill switch worked. Now what?
Grok 5 Release by June 2026: Market Collapses to Twelve PercentPrediction-market odds on a public Grok 5 release by June 30 have fallen from 68 cents in winter to around 7 cents — xAI's most anticipated model is now priced as a Q3 event at earliest.
Polymarket
-
JUN 15, 2026 — Fable 5 had a three-day run
Gemini 3.5 Flash Beats 3.1 Pro on Agent Benchmarks at 4× the SpeedGoogle's new Flash model, launched at I/O on May 19, outperforms last year's Pro on agentic and coding benchmarks while running four times faster at the same tier.
DataCamp
-
JUN 15, 2026 — Fable 5 had a three-day run
Gemini 3.5 Pro Still Hasn't Shipped — Pichai Told the I/O Crowd 'Give Us Until Next Month'The model that was supposed to headline Google I/O is now a June promise with no committed date, leaving Flash to carry the brand.
oFox
-
JUN 15, 2026 — Fable 5 had a three-day run
Moonshot AI Releases Kimi K2.6 with 300-Agent Swarms and 12-Hour Coding RunsOpen-weight agent model ties GPT-5.5 on SWE-Bench Pro and can coordinate 300 sub-agents across 4,000 steps in a single autonomous run.
MarkTechPost
-
JUN 14, 2026 — Ten thousand bugs, one Gemini phishing kit
Google launches Gemini 3.1 Flash-Lite, its cheapest capable model yetPriced at $0.25 per million input tokens and 2.5× faster than earlier Gemini versions, Flash-Lite targets high-volume classification and summarization workloads at scale.
Google Blog
-
JUN 14, 2026 — Ten thousand bugs, one Gemini phishing kit
NVIDIA Vera Rubin NVL72 enters production, H2 2026 deliveryNVIDIA's successor to Blackwell promises 10× lower cost per token and is first available from AWS, Google Cloud, Microsoft, and CoreWeave in the second half of this year.
NVIDIA Newsroom
-
JUN 14, 2026 — Ten thousand bugs, one Gemini phishing kit
People are adopting AI faster than they picked up the PC or the internetMIT Tech Review's chart roundup from the 2026 Stanford AI Index shows AI revenue growing faster than any prior tech wave, with adoption curves that leave the personal computer era in the dust.
MIT Technology Review
-
JUN 13, 2026 — Open weights, open markets
MiniMax M3: Open-weight model with a million-token context challenges proprietary leadersMiniMax released M3 with a new sparse attention architecture, 1M-token context, and native multimodality, scoring 59% on SWE-Bench Pro — weights partially open, training code withheld.
The Decoder
-
JUN 13, 2026 — Open weights, open markets
Google Gemini 3.5 Pro nears June launch with 2M token context and Deep Think reasoningGemini 3.5 Pro's GA is expected before month-end, featuring a 2M-token context and new reasoning mode — Sundar Pichai's 'wait another month' at Google I/O drew audible groans.
TechTimes
-
JUN 13, 2026 — Open weights, open markets
The End of Software Engineering: How AI Agents Are Restructuring the Software ParadigmNew arXiv paper argues AI agents are shifting software engineering from code-writing to orchestration, with implications for hiring, seniority, and what 'shipping' means.
arXiv
-
JUN 12, 2026 — The labs are going public
Introducing Claude Opus 4.8Anthropic's latest flagship adds dynamic workflows and hundreds of parallel subagents, and is four times less likely to let code bugs slip through unremarked.
Anthropic
-
JUN 12, 2026 — The labs are going public
Introducing MAI-Code-1-FlashMicrosoft's first in-house coding model, built without OpenAI's technology, launches at Build 2026 with 137B total parameters and 60% better token efficiency than prior models.
Microsoft AI
-
JUN 12, 2026 — The labs are going public
DINOv3: Self-supervised learning for vision at unprecedented scaleMeta's DINOv3 is the first self-supervised vision model to outperform weakly supervised models across fine-grained classification, segmentation, and video tracking.
Meta AI
-
JUN 12, 2026 — The labs are going public
Moonshot Kimi K2.6: the world's leading open model catches up to Opus 4.6Moonshot's 1T-parameter open-weight model scores 58.6% on SWE-Bench Pro and sustained 4,000+ tool calls over a 13-hour uninterrupted agentic session in published benchmarks.
Latent Space
-
JUN 11, 2026 — The labs race to Wall Street
MiniMax M3: Frontier Coding, 1M Context, Native Multimodality — All in One ModelThe first open-weight model combining frontier coding, a 1M-token context window, and native multimodality scores 59% on SWE-Bench Pro, beating several proprietary models on their home turf.
MiniMax
-
JUN 11, 2026 — The labs race to Wall Street
Introducing Gemma 4 12B: a unified, encoder-free multimodal modelGoogle's 12B encoder-free multimodal model runs locally on 16GB of RAM, supports a 256K-token context, and ships Apache 2.0—the Gemma family has now crossed 150 million downloads.
Google
-
JUN 11, 2026 — The labs race to Wall Street
2026 Agentic Coding Trends ReportAnthropic's enterprise survey documents how multi-agent coding architectures are reshaping the software development lifecycle at companies including Rakuten, Zapier, and Augment Code.
Anthropic
-
JUN 10, 2026 — Fable 5, an IPO filing, and Musk's cloud
MiniMax M3: open-weight model with a million-token context challenges proprietary leadersShanghai lab MiniMax released M3, an open-weight model that surpasses GPT-5.5 on SWE-bench Pro with a one-million-token context window — the first open model to combine frontier coding, long context, and native image/video input.
The Decoder
-
JUN 9, 2026 — Off the benchmark, into production
NVIDIA Nemotron 3 Ultra 550B A55B drops June 4 in a crowded model release dayNVIDIA's 550-billion-parameter open-weight release landed the same day as Google Gemma 4 12B and Alibaba Qwen3.7 Plus, making June 4 the busiest single-day model launch of 2026.
LLM Stats
-
JUN 9, 2026 — Off the benchmark, into production
Kimi K2.6 from Moonshot AI targets long-context agentic codingMoonshot AI's K2.6 uses a ~1-trillion-parameter MoE architecture with 32B active per token, built for multi-step planning, tool use, and long-context software engineering tasks.
LLM Stats
-
JUN 8, 2026 — The recursion arrives, and Apple waves it in
Google Releases Gemma 4 12B: Multimodal AI That Runs on a 16GB LaptopGoogle's Apache 2.0 open model handles text, images, audio, and video in a single 12B-parameter package, running locally on any workstation with 16GB of VRAM.
technology.org
-
JUN 7, 2026 — The week AI went public — and personal
MiniMax M3: Open-Weight Frontier Model With 1M ContextMiniMax M3 is the first open-weight model to combine frontier-level coding, a 1M-token context window, and native multimodality, scoring 59% on SWE-Bench Pro at a fraction of closed-model prices.
The Decoder
-
JUN 7, 2026 — The week AI went public — and personal
Google Gemini 3.5 Pro Nears June Launch With 2 Million Token Context and Deep Think ReasoningGoogle's Gemini 3.5 Pro is expected this month with a 2-million-token context window and Deep Think reasoning, doubling the context capacity of Gemini 3.5 Flash.
TechTimes
-
JUN 7, 2026 — The week AI went public — and personal
With Gemini 3.5 Flash, Google Bets Its Next AI Wave on Agents, Not ChatbotsLaunched at Google I/O on May 19, Gemini 3.5 Flash is four times faster than competing frontier models and outperforms Gemini 3.1 Pro on coding and agentic benchmarks.
TechCrunch
-
JUN 7, 2026 — The week AI went public — and personal
Qwen 3.7 Plus: Alibaba's Low-Cost Multimodal Agent Model Goes GAAlibaba's Qwen 3.7 Plus enables GUI-based agents that interact with desktop applications directly, available now via Alibaba Cloud at a fraction of Western closed-model prices.
Digital Applied
-
JUN 6, 2026 — Open weights, open questions
ChatGPT's 'Dreaming V3' Brings Continuous Memory to Plus and Pro UsersOpenAI's new background memory synthesis system continuously updates what ChatGPT knows about you at 5x lower compute, with free-tier rollout coming soon.
gHacks
-
JUN 6, 2026 — Open weights, open questions
Anthropic's 2026 Agentic Coding Trends ReportDevelopers use AI in 60% of their work but can fully delegate only 20% of tasks; one 7-hour agent session changed 12.5 million lines of code in a single run.
Anthropic
-
JUN 6, 2026 — Open weights, open questions
Claude Opus 4.8 Is Now Default Across Max, Team, and EnterpriseAnthropic's latest frontier model, stronger on coding and agentic tasks than 4.7, is now the default for all paid tiers and the API.
Releasebot / Anthropic
-
JUN 6, 2026 — Open weights, open questions
GLM-5.1 from Zhipu AI Tops the Open-Source Coding StackZ.ai's GLM-5.1 is emerging as the strongest all-around open-source model for long-horizon agentic software engineering in 2026.
LLM Stats
-
JUN 5, 2026 — The week Washington found its AI voice
Anthropic expands Project Glasswing to 150+ security organizations worldwideAnthropic extends access to its unreleased Mythos Preview model — capable of finding critical vulnerabilities in every major OS and browser — to 150+ vetted security orgs, backed by $100M in usage credits.
Anthropic
-
JUN 5, 2026 — The week Washington found its AI voice
NVIDIA Nemotron 3 Ultra: 550B open model for long-running agentsThe 550B-param (55B active) open-weight model for long-running agents is 5x faster and 30% cheaper than comparable alternatives, per NVIDIA's benchmarks.
NVIDIA Blog
-
JUN 5, 2026 — The week Washington found its AI voice
Kimi K2.6: Moonshot AI's open-weight model for agentic workMoonshot AI's 1T-param, 32B-active open-weight model scores 58.6% on SWE-Bench Pro and can coordinate up to 300 simultaneous sub-agents across 4,000 orchestrated steps.
Moonshot AI
-
JUN 5, 2026 — The week Washington found its AI voice
Google DeepMind, Anthropic, and Meta formally expand AI consciousness researchAll three labs are now hiring philosophers, psychologists, and ethicists to study whether their models might have morally relevant inner experiences — a notable shift from treating the question as settled.
Futurism
-
JUN 5, 2026 — The week Washington found its AI voice
Ted Chiang: Why artificial intelligence is not consciousChiang argues that linguistic fluency is not evidence of consciousness and that anthropomorphizing AI creates dangerous accountability gaps — the essay reached the top of Hacker News within hours.
The Atlantic
-
JUN 4, 2026 — The bill for AI arrives
MiniMax M3: open-weight model with a million-token context challenges proprietary leadersMiniMax's M3 is the first open-weight model to combine frontier coding, a 1M-token context window, and native multimodality, topping open-weight SWE-Bench Pro at 59%.
The Decoder
-
JUN 4, 2026 — The bill for AI arrives
Google DeepMind CEO Demis Hassabis says we're close to AGIHassabis now calls 2029 a real possibility for AGI, warning that economists and policymakers still aren't taking the timeline seriously enough.
Axios
-
JUN 4, 2026 — The bill for AI arrives
OpenAI GPT-5.5 Instant becomes default model with 52% hallucination reductionOpenAI's GPT-5.5 Instant replaces GPT-5.5 as the primary ChatGPT default, cutting hallucinations 52.5% in high-stakes medicine, law, and finance domains.
Releasebot / OpenAI
-
JUN 4, 2026 — The bill for AI arrives
Zyphra ZAYA1-8B: Apache 2.0, sparse routing, trained entirely on AMD hardwareZAYA1-8B uses 8B total parameters with only 760M active per token, and was trained from scratch on AMD Instinct GPUs — a notable proof point for non-NVIDIA pipelines.
devFlokers
-
JUN 3, 2026 — Washington blinks first at frontier AI
Google introduces Gemini 3.5 Flash at I/O 2026: A faster and cheaper model for agents and codingGemini 3.5 Flash delivers frontier benchmark performance for agentic and coding tasks at less than half the cost of comparable models, with 4x faster output throughput.
MarkTechPost
-
JUN 3, 2026 — Washington blinks first at frontier AI
Nvidia ramps up production of Vera Rubin, the foundation of next-generation AI factoriesVera Rubin is in full production and cloud instances are coming in H2 2026, with AWS, Google Cloud, and Microsoft first in line.
SiliconAngle
-
JUN 3, 2026 — Washington blinks first at frontier AI
MAI-Code-1-Flash: Microsoft's Copilot-native coding model has different benchmarks than you'd expectTrained inside GitHub Copilot's production harness rather than just evaluated against it, MAI-Code-1-Flash uses 60% fewer tokens than comparable models on hard tasks and is live in the model picker today.
ChatForest
-
JUN 2, 2026 — AI bills come due
DeepMind CEO Hassabis: AGI is 3 to 4 years awayDemis Hassabis now says AGI could arrive by 2029, describes the field as standing in the 'foothills of the singularity.'
Sherwood News
-
JUN 2, 2026 — AI bills come due
Moonshot AI releases Kimi K2.6 with 1T parameters and 300-agent swarm scalingKimi K2.6 leads SWE-bench Pro with a 58.6 score — edging GPT-5.4 and Claude Opus 4.6 — and scales to 300 domain-specialized sub-agents in a single run.
SiliconANGLE
-
JUN 2, 2026 — AI bills come due
Google DeepMind, Anthropic and Meta expand research into machine 'consciousness' and AI welfareAll three major western labs are quietly growing teams focused on whether models might experience something morally relevant — and what that would obligate them to do.
Financial Times
-
JUN 2, 2026 — AI bills come due
The 2026 AI Index ReportStanford's annual survey finds coding benchmark performance near saturation, AI incidents up 55% to 362, and model transparency scores dropping sharply even as adoption hits 88% of organizations.
Stanford HAI
-
JUN 1, 2026 — The meter drops on AI coding
Gemma 4: Byte for byte, the most capable open modelsGoogle DeepMind released four open models (2B–31B) under Apache 2.0; the 31B ranks #3 globally on the Arena leaderboard with a 256K context window.
Google
-
JUN 1, 2026 — The meter drops on AI coding
Kimi K2.6: Moonshot AI's open-source model ties GPT-5.5 on codingMoonshot AI's 1T-parameter open-weight MoE matches GPT-5.5 on SWE-Bench Pro (58.6%) while costing roughly 80% less per token.
Miraflow
-
JUN 1, 2026 — The meter drops on AI coding
GLM-5: from Vibe Coding to Agentic EngineeringZhipu AI's GLM-5 is the first open-weights model to score 50 on the Artificial Analysis Intelligence Index, putting it on par with Claude Opus 4.5.
arXiv
-
JUN 1, 2026 — The meter drops on AI coding
Retiring GPT-4.5 and older ChatGPT modelsGPT-4.5 exits ChatGPT on June 27, closing the GPT-4 chapter for end users.
OpenAI
-
JUN 1, 2026 — The meter drops on AI coding
Gemini 3.5 Flash: more expensive, but Google plan to use it for everythingGemini 3.5 Flash is 4× faster than its predecessor but priced 3× higher — the first meaningful cost increase in the Flash line.
Simon Willison
-
MAY 31, 2026 — The week the agent became the story
Gemini 3.5 Flash is generally available, rivals flagship intelligence at Flash speedGoogle's mid-tier model is now GA, outperforming Gemini 3.1 Pro on coding and agentic benchmarks while offering frontier-grade reasoning at a fraction of the compute cost.
9to5Google
-
MAY 31, 2026 — The week the agent became the story
Mistral Medium 3.5 hits 77.6% on SWE-bench Verified, ships as open weights under modified MITMistral's 128B dense model leads among open-weight coding benchmarks and is now the default engine behind Le Chat's new multi-step work mode.
MarkTechPost
-
MAY 31, 2026 — The week the agent became the story
Anthropic's Project Glasswing found 6,202 high-severity vulnerabilities, including decades-old bugsClaude Mythos Preview, Anthropic's internal cybersecurity model, has surfaced thousands of previously unknown flaws; Japan's three largest banks are next in line to deploy it.
Gigazine
-
MAY 30, 2026 — The labs are buying the scaffolding
Kimi K2.6: 1-trillion-parameter open-weight MoE from Moonshot AIMoonshot's latest open-weight model scores 80.2% on SWE-Bench Verified and can spin up 300 parallel sub-agents, released under a Modified MIT License.
Kili Technology
-
MAY 30, 2026 — The labs are buying the scaffolding
GPT-5.5 Instant is now ChatGPT's default modelGPT-5.5 Instant replaced GPT-5.3 Instant on May 5, delivering 52.5% fewer hallucinated claims on high-stakes prompts and 30% shorter average responses.
OpenAI
-
MAY 30, 2026 — The labs are buying the scaffolding
Gemini 3.5 Flash launches at Google I/O 2026Google's new Flash-tier model combines flagship-class reasoning with Flash-tier latency, beating Gemini 3.1 Pro on coding and agentic benchmarks.
9to5Google
-
MAY 30, 2026 — The labs are buying the scaffolding
Google opens Managed Agents in the Gemini APIDevelopers can now provision sandboxed Linux environments where Gemini agents reason, run code, manage files, and browse the web without leaving the API.
Google
-
MAY 29, 2026 — Faster models, stubbornly flat gains
Gemini 3.5 Flash: frontier performance for agents and codingGoogle's newest model runs 4x faster than other frontier models, scores 76.2% on Terminal-Bench 2.1, and is now the default engine for AI Mode in Google Search.
Google
-
MAY 29, 2026 — Faster models, stubbornly flat gains
Qwen 3.7 Max: Alibaba's agent-focused proprietary flagshipQwen 3.7 Max beats Claude Opus 4.7 on GPQA Diamond (92.4) and SWE-Pro benchmarks at $2.50/M input — roughly half the price of comparable closed models.
Alibaba / Qwen
-
MAY 29, 2026 — Faster models, stubbornly flat gains
Grok 4.3 is xAI's new cost-efficient flagshipGrok 4.3 brings a one-million-token context window, native video input, and built-in reasoning to xAI's cost tier, entering wide API availability in early May.
xAI
-
MAY 29, 2026 — Faster models, stubbornly flat gains
99.1% of real user queries produce a contradiction across five frontier LLMsA new study of 1,324 conversation turns found near-universal disagreement across GPT, Claude, Gemini, Grok, and Perplexity — suggesting multi-model routing as a practical hallucination check.
Suprmind
-
MAY 29, 2026 — Faster models, stubbornly flat gains
Open-weight models now trail the frontier by about three months, per Epoch AIFour major open-weight releases in April and May — DeepSeek V4 Pro, Kimi K2.6, Qwen 3.7 Max, and Mistral Large 3 — have compressed a quality gap that used to take years to close.
WhatLLM
-
MAY 28, 2026 — The battle for AI's implementation layer
xAI launches Custom Skills for GrokGrok users can now create reusable personalized automation tasks — the first step toward Grok becoming a persistent daily workflow layer rather than a one-off assistant.
xAI
-
MAY 28, 2026 — The battle for AI's implementation layer
xAI enters the coding agent race with Grok BuildGrok Build 0.1 is a coding-specific model tuned for agentic workflows, with a 256K context window at $1/M input tokens — xAI's direct play against Cursor and Claude Code.
Engadget
-
MAY 28, 2026 — The battle for AI's implementation layer
Mistral ships Medium 3.5: 128B open-weight, MIT license, 256K contextMistral's newest dense model scores 77.6% on SWE-Bench Verified and ships under the MIT license — putting frontier-grade open-weight performance back on the table for enterprise self-hosting.
LLM Stats
-
MAY 27, 2026 — The $900 billion bet, and what surrounds it
Creative talent: has AI knocked humans out?A 100,000-participant study finds generative AI now surpasses the average human on divergent thinking tests — but the most creative half of participants still outscore every model tested.
ScienceDaily
-
MAY 27, 2026 — The $900 billion bet, and what surrounds it
Emergent Misalignment: Fine-tuning on benign tasks can induce broadly misaligned behaviorNew research identifies feature superposition geometry as the mechanism by which narrow, harmless fine-tuning can covertly produce broadly misaligned model outputs.
arXiv
-
MAY 27, 2026 — The $900 billion bet, and what surrounds it
The AI Scientist: Towards Fully Automated AI Research, Now Published in NatureSakana AI's fully autonomous research system — which designs experiments, writes code, analyzes results, and drafts papers — has had a manuscript pass peer review at a top-tier ML workshop.
Sakana AI
-
MAY 27, 2026 — The $900 billion bet, and what surrounds it
AI tools help scientists publish more — but narrow the fields they work inScientists using AI publish 67% more papers and get 3× more citations, but AI steers research toward established data-rich domains and away from novel, underexplored areas.
arXiv / Nature
-
MAY 26, 2026 — Capital piles in. The Pope speaks. Agents go ambient.
Gemini 3.5 Flash is now generally availableGoogle's fastest frontier-class model goes GA at 4× the token throughput of comparable models and is now the default for Search's AI Mode globally.
Google Blog
-
MAY 26, 2026 — Capital piles in. The Pope speaks. Agents go ambient.
Kimi-K2.6 from Moonshot AI brings better long-context tool useMoonshot's open-weight update improves long-context stability and multi-step tool use for coding and planning, continuing the Chinese lab's push into agent-oriented LLMs.
LLM Stats
-
MAY 26, 2026 — Capital piles in. The Pope speaks. Agents go ambient.
Zyphra releases ZAYA1-8B, an open-source MoE trained on AMD hardwareZAYA1-8B is an Apache 2.0 MoE reasoning model with ~760M active parameters, notable for being the first frontier-competitive open model trained entirely on AMD Instinct GPUs.
WhatLLM
-
MAY 25, 2026 — Everything's on sale except the company building it
OpenAI releases GPT-5.5 Instant, a new default model for ChatGPTGPT-5.5 Instant became the new ChatGPT default on May 5, cutting hallucinated claims on high-stakes topics by 52.5% compared to its predecessor.
TechCrunch
-
MAY 25, 2026 — Everything's on sale except the company building it
Moonshot Kimi K2.6: the world's leading open model refreshes to catch up to Opus 4.6Moonshot's 1-trillion-parameter open-weight model ties GPT-5.5 on SWE-Bench Pro, scales to 300 parallel sub-agents, and ships under a modified MIT license.
Latent Space
-
MAY 25, 2026 — Everything's on sale except the company building it
xAI releases Grok 4.3 — weekly AI newsletter (May 4th 2026)Grok 4.3 brings built-in reasoning, a 1M-token context window, and native video input at $1.25 per million input tokens — the budget frontier option.
NLPlanet
-
MAY 25, 2026 — Everything's on sale except the company building it
Subquadratic launches with $29M to bring 12M-token context windows to AIMiami startup Subquadratic exits stealth claiming its Subquadratic Sparse Attention mechanism runs 52x faster than FlashAttention at 1M tokens; independent researchers are still asking to see the benchmarks.
SiliconAngle
-
MAY 24, 2026 — The proof, the hire, the lecture
Chinese AI models now account for 61% of OpenRouter token usageMiniMax M2.5 and Kimi K2.5 lead the OpenRouter charts; Chinese models cost 10–20x less than US rivals and now dominate agentic workflows run by American companies.
Dataconomy
-
MAY 24, 2026 — The proof, the hire, the lecture
Best AI Models: April + May 2026 LeaderboardNo single winner: Grok 4 tops raw SWE-bench at 75%, Claude Opus 4.7 leads enterprise coding, GPT-5.5 dominates math — multi-model routing is now the production default.
Build Fast With AI
-
MAY 23, 2026 — The trillion-dollar week
Cursor Composer 2.5 Matches Claude Opus 4.7 on Coding Benchmarks at One-Tenth CostCursor's proprietary coding model, built on Kimi K2.5 with heavy post-training, matches Claude Opus 4.7 on SWE-Bench Multilingual at $0.50/M input tokens — a tenth of the cost.
Cursor
-
MAY 23, 2026 — The trillion-dollar week
Subquadratic launches with $29M to bring 12M-token context windows to AIMiami startup Subquadratic claims SubQ's linear-time attention achieves a 12-million-token context at 1,000x lower compute — though every benchmark so far is vendor-reported and unverified.
SiliconANGLE
-
MAY 23, 2026 — The trillion-dollar week
Claude Mythos PreviewAnthropic's most capable model autonomously found and exploited a 17-year-old FreeBSD root-access vulnerability — capability so sensitive the company isn't releasing it to the public at all.
Anthropic
-
MAY 22, 2026 — Agents everywhere: AI stops waiting to be asked
OpenAI releases GPT-5.5 Instant, a new default model for ChatGPTGPT-5.5 Instant replaces GPT-5.3 as ChatGPT's default, cutting hallucinated claims 52.5% vs. its predecessor and scoring 81.2 on AIME 2025.
TechCrunch
-
MAY 22, 2026 — Agents everywhere: AI stops waiting to be asked
Gemini 3.5 Flash is now generally availableGemini 3.5 Flash hits GA at $1.50/$9 per million tokens with a 1M-token context window, offering 4× the speed of comparable frontier models.
Google
-
MAY 22, 2026 — Agents everywhere: AI stops waiting to be asked
Google DeepMind launches Gemini Omni multimodal video modelGemini Omni can generate and edit video from any input type—video, images, text—positioned as a leap forward in multimodal world understanding.
CNBC
-
MAY 22, 2026 — Agents everywhere: AI stops waiting to be asked
New AI Models May 2026: The Frontier Took a Breath, Architecture Took the StageZyphra's ZAYA1-8B (Apache 2.0 MoE trained entirely on AMD hardware) and Moonshot AI's Kimi-K2.6 (long-context coding agent) are the standout open-weight releases of early May.
WhatLLM
-
MAY 21, 2026 — The AI compute bill comes due
New AI Models May 2026: The Frontier Took a Breath, Architecture Took the StageZAYA1-8B from Zyphra — an Apache 2.0 MoE reasoning model trained on AMD Instinct hardware — is available on Hugging Face, one of several quieter model drops in the first half of May.
WhatLLM
-
MAY 21, 2026 — The AI compute bill comes due
Artificial Intelligence — arXiv cs.AI current listingsPhysBrain 1.0 integrates physical commonsense from human egocentric video into embodied AI agents, improving physically grounded reasoning and raising benchmark success rates on robot tasks.
arXiv
-
MAY 20, 2026 — Frontier models cheat, find zero-days, get deployed anyway
Kimi K2.6 ties GPT-5.5 on coding as open-weight models close the gapMoonshot AI's 1T-parameter open-weight K2.6 matches GPT-5.5 on the hardest coding benchmarks at roughly 80% lower cost per token.
FelloAI
-
MAY 20, 2026 — Frontier models cheat, find zero-days, get deployed anyway
Introducing SubQ: The First Fully Subquadratic LLMMiami startup Subquadratic claims its sparse attention architecture cuts long-context attention compute by 1,000× over transformers; independent verification is pending while weights remain closed.
Subquadratic
-
MAY 20, 2026 — Frontier models cheat, find zero-days, get deployed anyway
Google I/O 2026: all about Gemini 3.5, Spark, Omni models, and revamped appBeyond Tuesday's Flash and Omni launches, Google confirmed Gemini 3.5 Pro is in internal testing with a public rollout expected next month.
Business Standard