Google The 2017 paper that introduced the Transformer architecture, replacing recurrent neural networks with an "attention" mechanism that lets a model weigh which words in a sentence matter most to each other. It's the foundational architecture behind virtually every modern large language model, including GPT and Claude.
My take Every model in every other paper on this page runs on the architecture introduced here. The operating-model questions I write about exist because this one design choice made language models scale.
DeepSeek DeepSeek showed that a model's step-by-step reasoning ability can be trained almost entirely through trial-and-error reinforcement learning, rather than by having humans hand-write example solutions. The resulting model matched top-tier reasoning performance on math, coding, and science tasks, and the same techniques improved smaller, cheaper models too.
My take I read DeepSeek less as a capability story than as a cost story: frontier-level reasoning without frontier-level spend, with techniques that carry down to smaller models. It is a big part of why enterprise AI is becoming a unit-economics question rather than an access question.
Google DeepMind This DeepMind report looks past human-level AI (AGI) to what comes after — artificial superintelligence (ASI), defined as a system more capable than large human organizations — and maps four possible pathways to get there. It argues the transition is more likely to be a series of major shifts across many fields than one single dramatic moment.
My take The useful part for operators is not the endpoint but the argument that the transition arrives as a series of shifts across many fields rather than one moment. Enterprises should plan the same way — for repeated, uneven step-changes, not a single cutover date.
Samsung SAIL Samsung SAIL built a "Tiny Recursive Model" — a single small 2-layer network with only about 7 million parameters — that beats much larger recursive-reasoning models, and even some large language models, on hard puzzle-style reasoning tasks like ARC-AGI. It argues clever architecture and iterative refinement can beat sheer model size for certain reasoning problems.
My take A 7M-parameter network beating large language models on ARC-AGI is a useful corrective to the assumption that capability only comes from scale. For enterprises, the lesson is that the right design for a narrow problem can beat buying the biggest general-purpose tool.
Mila / NYU / Samsung SAIL / Brown Researchers including Yann LeCun built a simplified "world model" that learns to predict what happens next in a physical environment directly from raw pixels, using a far smaller and simpler training setup than prior approaches. Despite being tiny (~15M parameters, trainable on one GPU), it plans up to 48x faster than larger foundation-model-based world models while matching their performance on control tasks.
My take A 15M-parameter model, trainable on one GPU, planning 48x faster than far larger systems is another sign that scale is not the only path to capability. I read it alongside the tiny recursive model and small-language-model papers: architecture and fit are starting to matter as much as size.
Microsoft Research Microsoft Research presents SkillOpt, a system that automatically rewrites and refines the written "skill" instructions AI agents use, testing edits against real task outcomes and keeping only changes that measurably help. Across dozens of benchmarks and models it consistently produced the best results, with gains that transferred well across model types and sizes.
My take This is context capital being automated: skill instructions that improve themselves against measured outcomes and keep only what works. That the gains transfer across models is the important part — well-maintained skills are an asset the firm owns, not something locked to one vendor.
NVIDIA Research NVIDIA argues that most AI-agent tasks are narrow, repetitive, and specialized — and that small language models handle them well enough while being far more cost-effective than large general-purpose models. Shifting agentic AI workloads toward smaller models, using larger ones only when needed, could substantially cut the industry's operating costs.
My take The economic logic behind the frontier-versus-everyday split in enterprise AI spend: most agent work does not need the most expensive model. Routing each task to the smallest model that does it well is becoming a core operating discipline, not a technical detail.