Resources

Research worth your time.

Papers and reports I’ve read on where AI is heading and how it’s actually being used — each with a short audio overview and my take on why it matters.

AI Research

Model architecture, training, and capabilities.

Google

Attention Is All You Need

The 2017 paper that introduced the Transformer architecture, replacing recurrent neural networks with an "attention" mechanism that lets a model weigh which words in a sentence matter most to each other. It's the foundational architecture behind virtually every modern large language model, including GPT and Claude.

My take

Every model in every other paper on this page runs on the architecture introduced here. The operating-model questions I write about exist because this one design choice made language models scale.

Audio overview · 22 min
arXiv 1706.03762 Read the paper
DeepSeek

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek showed that a model's step-by-step reasoning ability can be trained almost entirely through trial-and-error reinforcement learning, rather than by having humans hand-write example solutions. The resulting model matched top-tier reasoning performance on math, coding, and science tasks, and the same techniques improved smaller, cheaper models too.

My take

I read DeepSeek less as a capability story than as a cost story: frontier-level reasoning without frontier-level spend, with techniques that carry down to smaller models. It is a big part of why enterprise AI is becoming a unit-economics question rather than an access question.

Audio overview · 15 min
arXiv 2501.12948 Read the paper
Google DeepMind

From AGI to ASI

This DeepMind report looks past human-level AI (AGI) to what comes after — artificial superintelligence (ASI), defined as a system more capable than large human organizations — and maps four possible pathways to get there. It argues the transition is more likely to be a series of major shifts across many fields than one single dramatic moment.

My take

The useful part for operators is not the endpoint but the argument that the transition arrives as a series of shifts across many fields rather than one moment. Enterprises should plan the same way — for repeated, uneven step-changes, not a single cutover date.

Audio overview · 27 min
arXiv 2606.12683 Read the paper
Samsung SAIL

Less is More: Recursive Reasoning with Tiny Networks

Samsung SAIL built a "Tiny Recursive Model" — a single small 2-layer network with only about 7 million parameters — that beats much larger recursive-reasoning models, and even some large language models, on hard puzzle-style reasoning tasks like ARC-AGI. It argues clever architecture and iterative refinement can beat sheer model size for certain reasoning problems.

My take

A 7M-parameter network beating large language models on ARC-AGI is a useful corrective to the assumption that capability only comes from scale. For enterprises, the lesson is that the right design for a narrow problem can beat buying the biggest general-purpose tool.

Audio overview · 16 min
arXiv 2510.04871 Read the paper
Mila / NYU / Samsung SAIL / Brown

LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

Researchers including Yann LeCun built a simplified "world model" that learns to predict what happens next in a physical environment directly from raw pixels, using a far smaller and simpler training setup than prior approaches. Despite being tiny (~15M parameters, trainable on one GPU), it plans up to 48x faster than larger foundation-model-based world models while matching their performance on control tasks.

My take

A 15M-parameter model, trainable on one GPU, planning 48x faster than far larger systems is another sign that scale is not the only path to capability. I read it alongside the tiny recursive model and small-language-model papers: architecture and fit are starting to matter as much as size.

Audio overview · 23 min
arXiv 2603.19312 Read the paper
Microsoft Research

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Microsoft Research presents SkillOpt, a system that automatically rewrites and refines the written "skill" instructions AI agents use, testing edits against real task outcomes and keeping only changes that measurably help. Across dozens of benchmarks and models it consistently produced the best results, with gains that transferred well across model types and sizes.

My take

This is context capital being automated: skill instructions that improve themselves against measured outcomes and keep only what works. That the gains transfer across models is the important part — well-maintained skills are an asset the firm owns, not something locked to one vendor.

Audio overview · 23 min
arXiv 2605.23904 Read the paper
NVIDIA Research

Small Language Models are the Future of Agentic AI

NVIDIA argues that most AI-agent tasks are narrow, repetitive, and specialized — and that small language models handle them well enough while being far more cost-effective than large general-purpose models. Shifting agentic AI workloads toward smaller models, using larger ones only when needed, could substantially cut the industry's operating costs.

My take

The economic logic behind the frontier-versus-everyday split in enterprise AI spend: most agent work does not need the most expensive model. Routing each task to the smallest model that does it well is becoming a core operating discipline, not a technical detail.

Audio overview · 17 min
arXiv 2506.02153 Read the paper

How people and organizations actually use AI.

Chinese Academy of Sciences et al.

A Survey of Vibe Coding with Large Language Models

This survey of over 1,000 papers examines "vibe coding" — developers directing AI coding agents and judging results by whether they work rather than reviewing the code line-by-line. It finds real productivity gains are inconsistent and depend heavily on disciplined practices like context management and structured collaboration workflows, not just the AI model's raw capability.

My take

The headline finding — that gains depend on disciplined practices like context management and structured workflows, not raw model capability — is the operating-model argument in miniature. The same tool produces very different results depending on how the work around it is designed.

Audio overview · 21 min
arXiv 2510.12399 Read the paper
OpenAI

How People Use ChatGPT

This NBER working paper tracks ChatGPT usage from its 2022 launch through July 2025, by which point it reached roughly 10% of the world's adult population. Non-work use grew from 53% to over 70% of all messages, with "practical guidance," "information seeking," and "writing" the three dominant use cases — and its main economic value coming from decision support in knowledge-intensive jobs.

My take

The finding I would point an executive to is that the main economic value comes from decision support in knowledge work, not task automation. Better judgment does not show up on a timesheet, which is exactly why AI value is so hard to find on the income statement.

Audio overview · 15 min
NBER Working Paper 34255 Read the paper
Microsoft

It's About Time: The Temporal and Modal Dynamics of Copilot Usage

Microsoft studied 37.5 million anonymized Copilot conversations from January-September 2025 and found sharp differences by device and time of day — mobile use skews toward health questions around the clock, while desktop use follows a clear 9-to-5 work pattern. The takeaway: people have woven AI into both work routines and personal life, differently depending on device and moment.

My take

The 9-to-5 desktop pattern is the telling one: at work, AI still follows the shape of the existing workday rather than changing it. That is the copilot plateau showing up in usage data — AI fitted around the old operating model, not redesigning it.

Audio overview · 15 min
arXiv 2512.11879 Read the paper
Perplexity

The Adoption and Usage of AI Agents: Early Evidence from Perplexity

The first large-scale study of how people actually use general-purpose AI agents in the open web, based on hundreds of millions of interactions with Perplexity's Comet browser agent. Early adoption skews toward wealthier, more educated users in knowledge-intensive professions, with usage concentrated in productivity and learning tasks and mostly personal rather than professional use.

My take

Early agent adoption looks like every prior technology wave — concentrated among wealthier, more educated knowledge workers, and personal before it is professional. The gap between people using agents on their own and firms redesigning work around them is where the real value is still waiting.

Audio overview · 16 min
arXiv 2512.07828 Read the paper
OpenAI

The Shift to Agentic AI: Evidence from Codex

OpenAI analyzes real usage data from its Codex coding agent and finds explosive growth in agentic AI use — weekly active users grew more than 5x in the first half of 2026, and people increasingly delegate longer, more complex tasks. Heavy users increasingly run multiple agents in parallel and reuse saved "skill" instructions to handle complex work.

My take

The clearest evidence yet that the unit of work is moving from the prompt to the delegated task. Heavy users running agents in parallel and reusing saved skills are doing at individual scale what I argue enterprises need to do: turn usage into context capital.

Audio overview · 20 min
arXiv 2606.26959 Read the paper
Anthropic

Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations

Anthropic mapped over 4 million real Claude.ai conversations to official U.S. occupational task categories and found software development and writing account for nearly half of all AI usage, with about 36% of occupations using AI for at least a quarter of their tasks. Usage splits roughly 57% "augmentation" (AI assisting a human) versus 43% "automation" (AI substituting for a task).

My take

The 57/43 split between augmentation and automation is the number I'd watch over time. Augmentation is the copilot plateau in data form, and the ratio shifting toward automation is the signal that firms are starting to redesign work rather than just speed it up.

Audio overview · 10 min
arXiv 2503.04761 Read the paper