Ideas worth reading
Short notes on AI and product practice.


AI Transformation Has to Serve the Customer
AI impact for its own sake is a weak journey. Whether B2B or B2C, operators should keep every process and cultural change anchored in a better, faster customer experience.
Read article
Why Boards Want AI—and the Ground Floor Moves Differently
Many C-level and board members feel underprepared to govern AI. The gap is board composition, the lag from questions to budget and talent—and sometimes operators who run ahead anyway.
Read article
Three Pillars of an AI-First Organization
Box’s own “Box on Box” journey points past individual productivity: redesign work with AI in mind, give agents clean governed data, and bring every employee along—not leave them to nights and weekends.
Read article
When Companies Call Themselves “Advanced” at AI, What Are They Measuring?
In Box’s AI report, self-rated advanced companies jumped from 8% to 64%. That leap tracks a shift from “should we?” to “how?”—and the best firms score impact, not tool rollout alone.
Read article
What Are Large AI Models Really Competing On?
A plain-language reading of Moonshot’s GTC talk and the Kimi K2.5 technical report: scaling is no longer just about size.
Read article
From Chatting Well to Getting Things Done
What the Kimi K2.5 technical report and Moonshot’s GTC talk suggest about multimodal AI and agent teamwork — in everyday language.
Read article
Start With One AI “Pace Car” Workflow
For the next 90–180 days, pick one exciting workflow and an engaged team—not a galactic strategy that overhauls everything at once. Early wins set the tone; boiling the ocean stalls it.
Read article
Why Enterprise AI Is Moving to Open Source
A plain-language reading of Glean founder Arvind Jain on 20VC: cost and “good enough” capability—not old data fears—are pushing enterprises toward open models.
Read article
Don’t Aim to Replace Yourself with AI
From Glean founder Arvind Jain’s 20VC interview: “replace yourself with AI” is the wrong goal. When everyone has the same tools, winners raise the bar—and strong companies may get larger.
Read article
Where Do Large Models Store Their Facts?
Attention lets tokens talk. Feed-forward layers do much of the deep processing — and hold a surprising share of factual knowledge, plus the scaffolding that makes deep stacks trainable.
Read article
Same Skeleton: How GPT, Claude, Gemini, and LLaMA Differ
Next-token prediction, decoding knobs, post-training, and the converging modern Transformer stack — what actually separates familiar model names.
Read article
Why Large Models Miscount the R’s in “Strawberry”
Models do not see letters the way people do. They see token IDs — and that design choice explains a classic “counting” failure.
Read article
What Attention Actually Does Inside a Large Model
A plain-language tour of queries, keys, values, causal masking, multi-head attention, induction heads, and why long prompts get expensive.
Read article
How Much Can a Language Model Actually Memorize?
An ICML 2026 study puts a number on GPT-style capacity: about 3.6 bits per parameter in BF16 — and explains why that is not a hard-drive size.
Read article
Model Parameters Are Not a Hard Drive
From an ICML 2026 memory-capacity study: weights mix sample details, shared rules, and training traces — and after a phase point, memorization yields to generalization.
Read article
When Bot Traffic Overtakes Human Traffic Online
A plain-language reading of Cloudflare CEO Matthew Prince’s claim that, in the first half of 2026, automated traffic passed human traffic—and why the timeline kept moving forward.
Read article
Why an AI Agent Can Visit Thousands of Sites to Pick a Camera
From Cloudflare CEO Matthew Prince: classic crawlers are not what bent the traffic curve—AI agents that research exhaustively are.
Read article
Bots Don’t Click Ads: Why the Web’s Business Model Must Be Rewritten
Matthew Prince’s blunt constraint: if machines dominate traffic, the ad model that funded free content for nearly thirty years starts to fail—and pay-per-crawl is one proposed path out.
Read article
Will Agents Kill the Software UI?
A plain-language take on a16z’s “headless software” discussion: when agents become users, the interface stops being the only door into the product.
Read article
Why Systems Like SAP Still Will Not Die
From an a16z conversation on headless software: enterprise stickiness lives in customized business logic, not in pretty screens—and agents do not erase that overnight.
Read article
The More Capable the Agent, the Easier It Is to Lower Your Guard
From Nenad Tomasev on the DeepMind podcast: agents fail more as tasks get complex, and “trust must be earned” is not the same as skipping checks.
Read article
Webpages Can Trick Your Agent Too
From Nenad Tomasev on the DeepMind podcast: the open web is the agent’s environment—hidden instructions, cloaking, and why defense in depth matters.
Read article
When a Million Agents Think Alike
From Nenad Tomasev on the DeepMind podcast: shared models create cognitive monoculture—correlated decisions, correlated failures, and new collusion risks.
Read article
Why Mid-Tier Models Can Suddenly Run Agents
A plain-language reading of Anthropic’s Claude Sonnet 5 launch: flagship-level agent skills are moving into the mid-tier price band.
Read article
Why Effort Level Matters More Than the Model Name
Anthropic’s Sonnet 5 cost–performance curves make a practical point: you can tune how hard a model tries, instead of only shopping by flagship vs mid-tier.
Read article
Why Stronger Models Are Often Safer — and Still Not Interchangeable
Anthropic’s Sonnet 5 safety notes show a useful split: mid-tier models can get safer for everyday agent use, while high-risk cyber skills still sit far above them.
Read article
Why “Intelligence Is Getting Free” Misses the Point
From Fei-Fei Li and David Rogier’s Silicon Valley Girl conversation: today’s AI talk mostly means language intelligence—and calling all intelligence “nearly free” is irresponsible.
Read article
Two Kinds of Workers May Matter Most
From Fei-Fei Li and David Rogier: a “barbell” future of top experts and high-agency generalists—where middling skill is easiest for AI to flatten.
Read article
AI Will Remake Tutoring—Not the Purpose of School
From Fei-Fei Li and David Rogier: one-on-one tutoring gets dramatically cheaper with AI—but schools should still aim to form people, not police chatbots.
Read article
AI Is Becoming a Utility—But Users Aren’t Buying “Intelligence”
From Sam Altman’s Stanford roundtable: early electricity sold night lighting, not “power.” AI may become infrastructure like water and electricity—once we find the right everyday promise.
Read article
Enterprise AI’s Bottleneck Is No Longer Technology
BCG’s AI at Work survey of nearly 12,000 workers points past models and tools: the real drag is strategy, organization, and people.
Read article
Strategy Beats Tools—by About Twenty Points
BCG’s AI at Work data: a clear AI strategy with limited tools outperforms a full toolkit with no direction—and honeymoon energy does not last forever.
Read article
The Joy Paradox: Happier at Work—and Harder Work
BCG’s AI at Work survey finds rising job satisfaction and rising cognitive load at once—and the firms that create the most value are also where people feel best.
Read article
Where Does the Time AI Saves Actually Go?
BCG finds many frontline AI users save a workday a week—yet most get little guidance on where that time should go, so efficiency leaks away.
Read article
Everyone Has Heard of AI Agents. Governance Has Not Caught Up.
BCG’s AI at Work survey: agent awareness and workflow integration are rising fast, while rules for human–AI teams and accountability still lag.
Read article
AI’s Value Is Collective—Not a Private Secretary
From Michael I. Jordan on Machine Learning Street Talk: today’s AI aggregates input from billions and should serve billions. Chasing a always-on private secretary misses the bigger economic problem.
Read article
Fluent Text Is Not One Step from General Intelligence
From Michael I. Jordan on Machine Learning Street Talk: the “first-step fallacy” treats today’s impressive demos as proof that all-purpose intelligence is near. Fluent mapping from input to output is not system-level thought.
Read article
Do AIs Really Understand the World—or Only Imitate?
A plain-language take on a hard question behind AGI timelines: if today’s systems only copy patterns in data, scaling alone may never produce real understanding.
Read article
Why Reinforcement Learning Is Close—but Not Yet Enactive
Among AI paradigms, RL best matches learning through interaction—yet Sutton and Rafiee argue three gaps still keep it from enactive cognition.
Read article
AI Dark Output: Why AI Changes Everything—Except the Productivity Stats
Tech headlines say AI is remaking the world; official productivity barely budges. Like Solow’s computer paradox, much of AI’s real value may never show up in GDP.
Read article
New Dark Output: The Work Nobody Did Before AI Made It Cheap
When a literature review drops from about $2,000 to about $2, we don’t do the same amount of work—we do vastly more. That value barely shows up beyond token bills.
Read article
AI Did Not Suddenly Get Smarter — It Finally Got Reliable Enough
A popular reading of OpenAI post-training lead Yann Dubois’s remarks: the “overnight leap” many people felt was a reliability threshold being crossed, not a break in the intelligence curve.
Read article
Why This Generation Cares More About “Twice as Fast” Than Leaderboard Points
From Yann Dubois’s public framing of GPT-era progress: the proudest product wins may be shifting the whole latency–performance curve, not a single flashy skill.
Read article
Why AI Technical Debt Compounds
From Anthropic’s Founder’s Playbook: ordinary tech debt accrues gradually; AI-generated debt drifts without a shared design model—and often breaks only after real users arrive.
Read article
Agent Bills Are Mostly for Reading Context, Not Writing Answers
A 2026 multi-institution study shows agentic coding costs are dominated by input tokens—often about 154 inputs for every 1 output.
Read article
Same Bug, Same Model—The Bill Can Still Swing Thirtyfold
Across SWE-bench Verified agent runs, token spend swings wildly between tasks—and even between repeats of the same task.
Read article
Whether a Model Saves Tokens Is a Habit—Not How Hard the Task Looks
On the same OpenHands + SWE-bench Verified setup, token efficiency gaps between frontier models stay large even when every model succeeds—or every model fails.
Read article
If It Feels Easy to a Human, That Still Won’t Predict the Agent’s Bill
Human time-to-fix ratings barely track agent token spend—and frontier models also systematically underestimate their own costs.
Read article
What a Large Model Is Thinking Can Finally Be Read in Words
Anthropic’s Natural Language Autoencoders turn high-dimensional activations into plain text—no labels required—so interpretability feels more like reading than decoding feature lists.
Read article
When a Model Knows It Is Being Tested—but Never Says So
Unspoken evaluation awareness is a core safety pain point. Anthropic’s NLA can surface it from activations—not only from what the model admits in text.
Read article
The Rhyme Was Planned Before the Line Ended
A poetry case study for Anthropic’s NLA: models plan rhymes ahead—and editing the explanation can causally change what they write next.
Read article