2026 10

Things I Learned - 13 Sep 2026

This week, I learned: Everything I own, owned suggests that agentic reverse-engineering of firmware helps us learn: Features the devices expose Hidden functionalities, e.g. Shure MV7 microphone has a command shell. Dependencies, supply chains and attack surfaces Interesting components, e.g. RTOS webcam has small face tracking and gesture detection models Change behavior, e.g. don’t turn on indicator while recording So, it’s possible (even likely) that my TV, phone, laptop, camera, fridge, car, vacuum cleaning robot, bluetooth headphone, … can be hacked by a rogue AI-assisted firmware update. “Leaving things alone is an underrated engineering skill.” From Software drives people insane. Across over a thousand forecasts, agents lost to a simple exponential weighted moving average forecast. Paper: RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases. Maybe I should ask agents to get the latest data first, rather than directly asking them to forecast, since the latter fetched less recent material. Claude Code offers function hooks if you enable CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1. These let you introduce code into almost any part of the Claude Code workflow, meaning you can convert Claude Code into practically amy kind of agent. (Probably a bit of competition to Pi.) However, neither ChatGPT nor I could figure out a use case I would need this for. We need more imagination! The Antropic team provide Claude Tag a separate service account. That’s an interesting portable pattern: giving agents a separate Linux username, GitHub account, email ID, database user ID, etc. is a pattern we understand and know how to govern. FutureSearch.ai is a forecasting app. I’m not sure what model is behind it or how good it is, but it decomposes a forecast into measurable signals, predicts those, and synthesizes. That’s a useful approach. For example, I asked it: Will LLM model routers and model routing companies grow in popularity and review or shrink by Jan 2027?. It broke it up into 5 forecast questions and answered them roughly as: Will OpenRouter’s reported weekly LLM token processing volume exceed 45 trillion tokens/week (about 1.8x its August 2026 level of ~25 trillion tokens/week) by January 31, 2027? (Yes, 95% chance. It’s already high and growing fast.) Will OpenRouter announce a new equity funding round, or otherwise be credibly reported to have reached a valuation above $1.3 billion, between August 2026 and January 31, 2027? (Yes, 84% chance. There seems to be market interest.) Will at least one LLM model-routing competitor to OpenRouter (e.g., Martian, Not Diamond, Portkey, Unify AI, TrueFoundry) announce a new equity funding round of $20 million or more between August 2026 and January 31, 2027? (Yes, 68% chance. VCs will want to fund, and competitors exist.) Will a major AI lab or cloud provider (OpenAI, Google, Microsoft/Azure, Amazon/AWS, Anthropic, or Meta) launch or significantly expand, between August 2026 and January 31, 2027, a native product feature that automatically routes a given request among multiple materially different underlying LLMs based on cost, task, or quality? (Yes, 91% chance. Microsoft already has one; Google launched a preview; AWS will likely announce in re:Invent in Dec) Will Google Trends relative search interest (US, web search) for the term ‘LLM router’ be higher, on average, in December 2026 than it was in July 2026? (No, 25% chance. July 2026 was exceptionally high volume.) In An Alien Mind, Jakub Pachocki, Chief Scientist at OpenAI, was quite instructive. Here’s my takeaway: Models could keep growing smarter at the same speed. We can improve them where capability is measurable, like maths. In fuzzy areas, we’re not even sure how capable they are. Values are fuzzy. Making AI follow our values is tricky. We train models to follow their constitution. But they sometimes fail outside of their training examples. We feed models alignmed data. But when trained against hard objectives, they gently bend rules. We watch models’ thoughts. We avoid feedback on thoughts - so models won’t hide them. But models interact with agents & tools while thinking, so we need to supervise thoughts. Nowadays,models think without verbalizing. They manipulate their own reasoning. So we’re exploring confessions and monitoring internals. Still… best to tighten defenses. We’ll use AI to research how. Meeting people who have a target AND who control scarce resources is a great exercise in humility. Principals of elite private schools, partner managers of top software companies, any officer with a quota (police, income tax, bank loan, IT compliance), etc. You learn to grin while bearing the pain of being with them. Thanks to agents, it’s easy enough to maintain an Android and iOS mobile application separately #ForNow, rather than incur the overhead of React-Native (or other cross-platform frameworks). Shopify is making testing easy by “… designing our app architecture to work for both humans and agents.” Use re.prefixmatch() instead of re.match() in Python 3.15+. This article captures the reason well. (I failed the quiz at the start despite almost 2 decades of Python programming - and LLM atrophy). You can run Linux distributions in the browser. For example, this is a simple, embeddable buildroot distribution that runs purely in the browser. There’s Nix. There’s Alpine Linux. Interestingly, curl https://example.com/ works on Alpine Linux, unconstrained by same-origin policies. It is relayed by the host (bellard.org in this case) via WebSockets, so it can even ssh into other servers. ChatGPT ChatGPT’s Cloud Browser doesn’t forward all events - so it gets stuck on captchas, like Cloudflare’s, when visiting sites like StackOverflow. Here’s an example. Several top-level domains have over 50% of new registrations in 2025 blocklisted. Scammers use new domains extensively. But policing new domains also stops genuine protesters, so it’s not clear what the right approach is. The purpose of DNS is to spread scams. Questions I was asked Week ending 13 Sep 2026 ...

My predictions in 2025

In 2025, I made a number of predictions on this blog. (Not intentionally. I just said stuff.) I asked ChatGPT to audit them. It selected 440 claims, filtered out vague or pending ones, and verified the rest. Here’s what I got right and wrong. 🟢 “My chat will overtake search in 12-18 months. When ChatGPT becomes my primary lens on knowledge…” (chatgpt-vs-google-usage.md:L27). Audit: This actually happened in your own browsing data: search led through April 2026; in May, chat jumped to 2,211 visits vs 1,319 search visits, and stayed comfortably ahead thereafter. 🟢 “Typed languages are better suited for vibe coding. This will likely lead to the growth of typed languages (TypeScript, Rust, Go) but also of typing in untyped languages (e.g. Python).” (things-i-learned-10-aug-2025.md:L33). Audit: TypeScript became GitHub’s #1 language in August 2025 and grew 66% YoY; GitHub itself explicitly connects the rise of typed languages to more reliable AI-assisted coding. This is unusually strong because you got both the direction and mechanism. (The GitHub Blog) 🟢 “Code Mode … is a smart way to use MCPs and a very likely future direction. Using LLMs to write code to call MCPs rather than directly.” (things-i-learned-05-oct-2025.md:L49). Audit: OpenAI’s current Responses API has essentially this as a named capability: Programmatic Tool Calling lets a model write and execute programs that coordinate multiple tools and intermediate results. (OpenAI) 🟢 “CLI optimization for LLMs will likely emerge. More CLIs (and wrappers / hooks in the shell) will improve output and error contexts for LLMs…” (things-i-learned-27-jul-2025.md:L86). Audit: By May 2026 you were yourself testing rtk, a CLI proxy explicitly producing compact agent-friendly command output; across 216 commands you measured about 50% token reduction. That is almost exactly the wrapper you predicted. 🟢 “In the future, AI that works directly with file systems, Model Context Protocols, and local APIs are likely to become more important.” (features-actually-used-in-an-llm-playground.md:L76). Audit: File-system-native coding agents became mainstream, while MCP expanded into both hosted and local integrations; Anthropic now packages local MCP servers as easy-to-install desktop extensions. (Claude Help Center) 🟢 “Agents are slow. Parallelizable tools … will grow. Tool speed … will become more important.” (things-i-learned-22-jun-2025.md:L106). Audit: Parallelism has become a central agent UX: OpenAI’s Codex app is explicitly built around managing multiple agents simultaneously; at the extreme, internal users now accumulate more than 60 hours of agent turns per day by parallel execution. (OpenAI) 🟢 “Companies-of-one will grow. Sole founder can handle support functions.” (things-i-learned-24-aug-2025.md:L14). Audit: Nasdaq’s Economic Institute finds one-person US business applications up more than 20% since early 2025, with essentially all recent application growth coming from solo businesses. Stripe separately reports solo founders reaching 63% of Atlas C-corps in Q2 2026. (Nasdaq) 🟢 “We will move towards an organization structure where developers are embedded with business teams rather than working as a separate group. Sort of like embedded executive assistance instead of a central typing pool.” (things-i-learned-08-jun-2025.md:L49). Audit: Forward-deployed engineer demand reportedly increased 42-fold from 2023-25, with roughly 9,000 roles globally by early 2026. The role is almost precisely “technical people embedded with the business to make AI work in its real environment.” (Reuters) 🟢 “Shadow apps will grow. Anyone can code. Users build apps with prompts, sheets, agents, outside of IT SDLC. Like Excel sheets.” (things-i-learned-24-aug-2025.md:L18). Audit: Microsoft now explicitly describes a “new wave of shadow AI”: users installing coding/desktop/SaaS agents outside traditional IT governance, and has built discovery products specifically for unmanaged AI applications and agents. (Microsoft) 🟡 “Agents generate diffs/PRs. Tools to edit and comment on these online will emerge.” (things-i-learned-22-jun-2025.md:L107). Audit: GitHub now measures PRs created and merged by Copilot coding agent, while review comments can be handed directly to the agent with “Fix with Copilot,” including batches of review feedback. That’s almost verbatim fulfillment. (The GitHub Blog) 🟡 “Models’ ability to orchestrate longer workflows will improve. Factor that into your application design.” (things-i-learned-10-aug-2025.md:L44). Audit: By mid-2026, OpenAI reports large increases in requests corresponding to >30-minute, >1-hour and even >8-hour human tasks, while Codex explicitly targets long-running tasks spanning hours or longer. (OpenAI) 🟡 “Code review process will be re-invented.” (things-i-learned-22-jun-2025.md:L109). Audit: GitHub has rebuilt Copilot review around an agentic architecture that gathers broader repository context, uses tools, produces findings, and can hand fixes to another coding agent. This is substantially more than autocomplete added to old review. (The GitHub Blog) 🟡 “Domain expertise will therefore become even more valuable in the near future.” (things-i-learned-20-apr-2025.md:L39). Audit: 2026 hiring evidence points toward domain/product expertise becoming more important rather than pure coding alone, particularly as AI handles more implementation and firms need people who can connect it to actual business functions. (Reuters) 🟡 “Validation is the New Bottleneck: Since coding is now much faster, the critical, time-consuming task has shifted to reviewing, testing, and validating the LLM’s output.” (things-i-learned-17-aug-2025.md:L90); you also predicted “The Quality Control (QC) function will become larger and more critical” (L95). Audit: GitHub has now productized exactly that bottleneck in Code Quality; more than 10,000 enterprises used its preview, and GitHub explicitly frames AI-accelerated code output as creating the need for trustworthy pre-merge quality validation. (The GitHub Blog) 🟡 “Agents generate technical debt faster than humans. Solving this will become a major problem/opportunity.” (things-i-learned-22-jun-2025.md:L114). Audit: GitHub’s 2026 Code Quality launch is close to a commercial instantiation of this forecast: AI increases code output, so automated quality/debt detection and remediation moves earlier into the development cycle. (The GitHub Blog) 🟡 “Governance will grow. Non-experts are acting like experts. Validation is more important.” (things-i-learned-24-aug-2025.md:L19). Audit: The companion to shadow AI has indeed been governance: Microsoft now ships specific discovery, monitoring and governance for unmanaged AI agents, while NIST has continued expanding formal GenAI evaluation tooling. (Microsoft Learn) 🟡 “Soon, we won’t just follow a lesson plan – we’ll have lessons built just for us. AI will track how we learn and adapt in real time. It’ll feel like having a personal coach in your back pocket.” (o3-is-now-my-personalized-learning-coach.md:L91). Audit: ChatGPT Study Mode now asks what the learner knows, adapts explanations, checks understanding, works from uploaded course material, and uses memory to personalize support; OpenAI explicitly describes the objective as personalized learning support available to any student. (OpenAI Help Center) 🟡 “Cost is going down so quickly right now that all you have to do is wait, and stuff will become available for a very affordable or even a free price.” (things-i-learned-16-mar-2025.md:L116). Audit: The broad direction held. OpenAI cut GPT-5.6 Luna API prices by 80% in July 2026 while simultaneously improving capability-per-dollar. The “all you have to do” part is hyperbole, but the price-curve forecast was right. (OpenAI) 🟡 relayed: “Control of chips and GPU compute is what will likely be the gameplay to control AI dominance globally.” (things-i-learned-02-feb-2025.md:L16, attributed there to Dario Amodei). Audit: Advanced-AI-chip export licensing remains an explicit geopolitical control mechanism in 2026, including restrictions and license review for H200/MI325X-class accelerators going to China. 🔴 “AI closes the gap between junior & senior devs – even when both use AI. Quality doesn’t suffer much. So onboarding can be faster, compensation ladder may shorten.” (things-i-learned-03-aug-2025.md:L52). Audit: The emerging evidence says AI changes the work but does not erase the expertise gap: experienced developers are better at steering/delegation, while low-experience AI-heavy contributions incur substantially more review and lower acceptance. (arXiv) 🔴 “LLMs already deliver hours of analyst work in minutes. Entry-level roles WILL vanish.” (goodbye-mba-hello-ai.md:L17). Audit: The labor-market warning was directionally good, but “vanish” is a major magnitude error. Stanford finds a meaningful relative decline among 22-25-year-olds in highly AI-exposed jobs, while employment remains substantial and overall exposure groups still show employment growth. “Entry-level hiring contracts sharply” would have scored much better. (Stanford Digital Economy Lab) 🔴 “Coders micro-manage LLMs. I think a novice will be more efficient and get better results than me.” (how-to-visualize-data-stories-with-ai-lessons.md:L285). Audit: Current empirical work points the other way in real software work. In a 22,953-PR study, lower-experience AI-heavy developers received 4.5* more review comments, had 31% lower acceptance, and took over 5* longer to resolve issues; qualitative work likewise finds experienced developers better at delegation and control. (arXiv) 🔴 relayed: “API access from model providers will shrink. Selling tokens is not a viable business model given lowering costs.” (things-i-learned-23-mar-2025.md:L19, from the Alexander Doria notes immediately above it). Audit: Almost exactly backwards. Model providers expanded their APIs into richer agent platforms, and token-metered API access remains a core commercial model - including premium pay-as-you-go modes. (OpenAI) 🔴 “APIs are likely to be replaced by just chat requests that will do the same thing. APIs might be replaced by RPA, where somebody uses a chatbot to do the equivalence instead.” (things-i-learned-16-mar-2025.md:L111-L112). Audit: Chat did become a front end, but the implementation moved toward more APIs underneath, not fewer: tool APIs, Responses, MCP, computer-use interfaces and programmatic tool calling are now the substrate agents use. (OpenAI) 🔴 “Software companies build ‘SaaS’-like apps today. Agents will replace apps. Instead of UI, workflows, and app logic, they’ll engineer prompts, APIs, and evals.” (agents-will-replace-saas-apps.md:L12). Audit: The interface-shift was right; “replace” was not. Gartner now forecasts agentic AI may expose roughly 20% of SaaS application spending by 2030 - meaning substantial disruption, not app extinction. Agents are often a new interaction layer over systems of record and APIs. (Gartner) 🔴 relayed: Models will “internalis[e] workflows … to wipe out the apps and workflow space.” (things-i-learned-23-mar-2025.md:L17, from Alexander Doria notes). Audit: “Internalize capabilities” was insightful; “wipe out” was the failed extrapolation. Enterprise applications remain a very large substrate even in Gartner’s fairly aggressive agentic-AI forecast. (Gartner) 🔴 “Demand for SaaS (one-size-fits-all) will shrink.” (things-i-learned-06-apr-2025.md:L89). Audit: Not yet. For example, Gartner forecasts Indian SaaS spending growing 18.9% in 2026, from $3.9B to $4.6B. AI is changing SaaS economics and seat licensing, but current demand is still growing rather than shrinking. (Gartner) 🔴 “The early majority have come in… Soon the late majority will come in asking for existing solutions that have already solved their problem for many others.” (things-i-learned-03-aug-2025.md:L44). Audit: This mapped your client/audience experience onto population adoption much too quickly. In 2026, US Census data put business AI usage around 17-20%, nowhere near a conventional late-majority phase. (Census.gov) 🔴- baseline error: “Given the cost and accessibility of drones, I guess drone terrorist attacks will soon emerge.” (things-i-learned-16-nov-2025.md:L43). Audit: They had already emerged. The UK government documented Daesh using small armed remotely piloted aircraft carrying grenades in Iraq in 2017; the UN had already been studying weaponized UAS use by non-state armed groups for terrorism-related purposes before this 2025 post. This is therefore a clean failure to establish the baseline, not a future hit. (GOV.UK) 🔴 relayed: “Personal writing with connection won’t go away. AI can’t give you heartbreak. But the rest of non fiction writing will vanish.” (things-i-learned-30-mar-2025.md:L69, under “Notes from Writing with AI”). Audit: Nonfiction is under real pressure, but “vanish” is nowhere close. UK nonfiction publishing still generated about GBP1.0B in 2025, down only 3%; the wider publishing industry reached record revenue. (Publishers Association) Legend: ...

The Falling Cost of Intelligence

It’s amazing to watch the cost of intelligence falling. In Nov 2023, we had college-junior level intelligence for $10 per million tokens, i.e. it would take them $10 to read and process something as large as all seven Harry Potter books. ...

How IMF mis-forecasts GDP growth

The IMF forecasts GDP growth every year. Their forecasts for the current year are slightly low. Their forecasts for the next year are slightly high. After that, it remains high. Some forecasts, like China, Singapore, UAE, Equatorial Guinea are consistently low. Other forecasts, like Japan, Congo, Mexico, Pakistan are consistently high. The interesting meta-pattern is how this sort of past-forecast analysis can be done for any topic. This emerged from an Ethan Mollick post and then I asked: ...

How AI bottlenecks shift

I wrote about my changing AI opinions. At least some of this is because the industry is moving so fast that the bottlenecks keep shifting. Here are four examples of how we AI couldn’t do something (the bottleneck), but that became possible, and the bottleneck shifted - changing the way we work. It’s good to keep this in mind when thinking about AI. Coding: “It can’t write useful code. We can’t get real help.” But in Sep 2022: GitHub finds Copilot developers are 55% faster. “It writes code but doesn’t know our codebase. We can’t let it touch real projects.” But in Feb 2024: Gemini 1.5 Pro has 1M-token context ~ 30K LOC". Cursor indexes code. “It understands the repo but can’t ship a fix on its own. We can’t hand it a whole issue.” But in Mar 2024: Devin solves 14% of SWE-bench - up from 2%.. Verified SWE-Bench is now 70%+. “It ships fixes, but we can’t review them fast enough or trust they’re stable.” Oct 2024: DORA 2024 finds AI hurt both throughput and stability. Now: Sep 2025: DORA 2025 finds is positive but stability stayed negative. Now: Jul 2025: METR’s RCT finds experienced devs 19% slower. Agents ...

Where Enterprise AI is headed

A podcast host sent me eight questions. Instead of rehearsing answers in my head, I used ChatGPT with Local MCP to read 6 months of call transcripts and find the best examples: Iteration 1: Here are questions I have been asked to answer in a podcast. Help me prepare with examples. For each question, go through my transcripts or emails and find examples relevant to the question and share (for each relevant example) a summary, how it’s relevant, and the relevant verbatim quotes from the transcript. Iteration 2: Mention WHO said it. Emphasize the most important parts. Do a second pass. More examples. Disprove your own hypotheses with evidence to the contrary and retain what remains robust. Iteration 3: Do a third pass. Find more real-life examples. Try and disprove yourself even harder. Share the best examples for what survives - not all. Same format. Iteration 4: Ensure diversity of client examples. For example, in Q2, all three are the same client. Extend to add / replace examples - ideally with better ones. Then I used Claude with examples of my writing style to summarize it in my voice. ...

Coding agents ARE the new software

Increasingly, I use coding agents instead of writing software. For example, I built a Blog UMAP. Then, I built Calvin UMAP. And more. But instead of building re-usable software, I just ran Claude with prior context. Increasingly, I use coding agents to run software. For example, I use Codex to classify my expense receipts. It writes re-usable code, but I run it using Codex, and it updates the code with new/edge cases. ...

The Future of Work with AI

I often research how the world will change with AI by asking AI. Today’s session was informative. I asked Claude, roughly Economics changes human behavior. As intelligence cost falls to zero, here are some changes in my behavior [I listed these]. Others will have experienced behavioral changes too. Search online and synthesize behavioral changes. It said this. 🟡 People spend time on problem framing & evaluation. AI can execute the middle. (I’m OK at this. Need to do more framing + evaluation.) 🟢 People don’t plan, they just build. (I’m prototyping a lot.) 🟢 People build personal data & context. (I’m mining my digital exhaust.) 🔴 People queue work for agents, delegating into the future. (I’m not. I need to do far more of this.) 🟢 People shift from searching to asking for answers. (I do this a lot, e.g. this post.) 🟡 People are AI-delegating junior jobs and developing senior level taste early. (Need to do more.) 🟡 People treat unresolved emotions as prompts. (Need to do more.) Rough legend: 🟢 = Stuff I know. 🟡 = I kind-of know. 🔴 = New learning. ...

AnalAIzing Cloud Costs

I have a GitHub Education since I teach at IITM. But if I switch back to a free account, how much would I need to pay? I asked Codex (5.3, xhigh): My GITHUB_TOKEN is in .env. Go through my GitHub billing. Ignore the $100 sponsorships I make. Other than that, my current metered usage is $6.71 for Feb 2026 (which is included in my billing plan). $0.35 comes from sanand0/exam and $0.34 from sanand0/blog and so on. That’s coming mostly from “Actions Linux”, occasionally “Actions Storage”. Pick a few of the top repos and tell me what I should do to make the cost zero - or reduce the cost as much as possible. See if there’s a pattern across repos. ...

When LLM prices fall 10x every year

In Feb 2024, Claude 3 Opus was the best model, at $15/MTok. In Jul 2024, GPT 4o Mini reached that quality at 10% of the price. In Dec 2024, DeepSeek v3 reached that quality at 1% of the price. Video See the interactive version If the price continues to fall 10x every 11-12 months or so (and it has been), then in a year, a Claude 4.6 Opus like model will cost 1/10th of the $5/MTok today, and in 2 years, 1/100th of that. ...

2006 1

Timeline of trends and events

Timeline of trends and events from 1750 to 2100 (yes, that’s next century).

2004 1

Exchanges using Wisdom of Crowds

Alternate exchanges: Hollywood Stock Exchange, Innovation futures, Blogshares. These exchanges are supposed to be remarkably accurate in predictions.

2002 1

More on expectations higher than expected

After reading my post on the ET article mentioning “expected to see a higher than expected rise”, a certain CA gold-medallist friend of mine wrote back this obscure note that I refuse to understand: … if you take it literally it is not possible. To put it more technically, something called a law of iterated expectation comes to play. Today’s expectation of tomorrow’s expectation about what will happen day after is just today’s expectation of what will happen day after. ...