<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>vibe-coding on S Anand</title>
    <link>https://www.s-anand.net/blog/tag/vibe-coding/</link>
    <description>Recent content in vibe-coding on S Anand</description>
    <generator>Hugo -- 0.164.0</generator>
    <language>en-us</language>
    <lastBuildDate>Fri, 05 Jun 2026 09:41:33 +0800</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/tag/vibe-coding/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>My changing AI opinions</title>
      <link>https://www.s-anand.net/blog/my-changing-ai-opinions/</link>
      <pubDate>Fri, 05 Jun 2026 09:41:33 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/my-changing-ai-opinions/</guid>
      <description>&lt;p&gt;I asked Claude about my AI opinions.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Based on my transcripts and blog posts, find the three claims I make most consistently, the three I&amp;rsquo;ve quietly reversed, and the one assumption I&amp;rsquo;ve never questioned but everything depends on.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;!-- https://claude.ai/chat/1d41b5a3-7963-4e87-aade-c44a37d5f355 --&gt;
&lt;p&gt;Here are things I&amp;rsquo;ve changed my opinion on:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;THEN&lt;/strong&gt;: One frontier model will win - not specialization. &lt;strong&gt;NOW&lt;/strong&gt;: Gemini for media, Claude for strategy/style, GPT for rigor. SLMs as tools.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;THEN&lt;/strong&gt;: Carefully curate my course content. &lt;strong&gt;NOW&lt;/strong&gt;: Give students prompts directly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;THEN&lt;/strong&gt;: Web apps are differentiated artifacts. &lt;strong&gt;NOW&lt;/strong&gt;: HTML is easier to generate than PPT - a signal of slop, not craft.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;THEN&lt;/strong&gt;: Human in the loop. &lt;strong&gt;NOW&lt;/strong&gt;: Human NOT in the loop, bottlenecking it. On-the-loop, etc. is fine.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;THEN&lt;/strong&gt;: Minimal single-agent loop, avoid sub-agents&amp;quot; &lt;strong&gt;NOW&lt;/strong&gt;: Multi-agent, sub-agent, and agent teams.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;THEN&lt;/strong&gt;: Avoid MCP, prefer SKILLS.md. &lt;strong&gt;NOW&lt;/strong&gt;: Use MCP because integrating with Claude / ChatGPT / &amp;hellip; is easy.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;There are the top contradictions in my opinions.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&amp;ldquo;Vibe code everything, end-to-end&amp;rdquo; &lt;em&gt;vs&lt;/em&gt; &amp;ldquo;Don&amp;rsquo;t commit based on vibe-coding, it&amp;rsquo;s not fully reliable&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Experience is a liability&amp;rdquo; &lt;em&gt;vs&lt;/em&gt; &amp;ldquo;Domain expertise is an edge&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Paid plans are fine for private data&amp;rdquo; &lt;em&gt;vs&lt;/em&gt; &amp;ldquo;Don&amp;rsquo;t assume paid = private&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Get out of the way / delegate everything&amp;rdquo; &lt;em&gt;vs&lt;/em&gt; &amp;ldquo;Verification is the whole job&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Build, don&amp;rsquo;t plan&amp;rdquo; &lt;em&gt;vs&lt;/em&gt; &amp;ldquo;Plan → Correct → Execute&amp;rdquo;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Some things, I have never questioned.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;For numbers, math, or correctness, make AI write and run code &amp;ndash; never trust prose arithmetic.&lt;/li&gt;
&lt;li&gt;Judge AI against human accuracy, never against perfection &amp;ndash; because experts disagree among themselves.&lt;/li&gt;
&lt;li&gt;Use AI heavily &amp;ndash; reach for it first, high volume.&lt;/li&gt;
&lt;li&gt;Don&amp;rsquo;t build a foundation model from scratch &amp;ndash; steer existing general models instead.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Things I usually say, &lt;strong&gt;but&lt;/strong&gt; there are exceptions.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&amp;ldquo;Always verify&amp;rdquo;. &lt;strong&gt;But&lt;/strong&gt; blindly trust AI in non-core areas where you are unskilled (such as personal finance).&lt;/li&gt;
&lt;li&gt;&amp;ldquo;AI is improving fast&amp;rdquo;. &lt;strong&gt;But&lt;/strong&gt; there is a jagged edge, to verify before upgrading.&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Don&amp;rsquo;t build models&amp;rdquo;. &lt;strong&gt;But&lt;/strong&gt; go ahead if it&amp;rsquo;s easy and beats LLMs clearly.&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Models keep getting cheaper&amp;rdquo;. &lt;strong&gt;But&lt;/strong&gt; not in May 2026.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-05-my-changing-ai-opinions.avif&#34;&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 29 Mar 2026</title>
      <link>https://www.s-anand.net/blog/things-i-learned-29-mar-2026/</link>
      <pubDate>Sun, 29 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-29-mar-2026/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://thekidshouldseethis.com/&#34;&gt;The Kids Should See This&lt;/a&gt; - great collection of videos for curious people. &lt;a href=&#34;https://thejeshgn.com/2026/03/27/weekly-notes-13-2026/&#34;&gt;Thej&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;A jury fined Meta and YouTube $4.2m and $1.8m for building addictive features in their products. That&amp;rsquo;s a first. &lt;a href=&#34;https://www.nytimes.com/2026/03/25/technology/social-media-trial-verdict.html&#34;&gt;NY Times&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;I think AI-type tools will actually revolutionize the experimental side of math, where you don’t care so much about individual problems and the process of solving them, but you want to gather large-scale data about what things work and what things don’t.&amp;rdquo; &lt;a href=&#34;https://www.dwarkesh.com/p/terence-tao#:~:text=gather%20large%2Dscale%20data&#34;&gt;Terence Tao&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The &lt;a href=&#34;https://en.wikipedia.org/wiki/Hedonic_treadmill&#34;&gt;hedonic treadmill&lt;/a&gt; (which roughly quantifies a Buddhist principle) says that we revert to a &lt;a href=&#34;https://en.wikipedia.org/wiki/Hedonic_treadmill#Happiness_set_point&#34;&gt;happiness set point&lt;/a&gt; (which varies by individual). Worse, those who experience a high kick (e.g. a lottery) don&amp;rsquo;t get enough kick from normal wins (contrast effect) &amp;ndash; &lt;a href=&#34;https://gemini.google.com/share/9e8a904b34bb&#34;&gt;Interactive explainer&lt;/a&gt;. &lt;!-- https://gemini.google.com/app/b676e7571e5cbc85 --&gt; The happiness neutral&lt;/li&gt;
&lt;li&gt;As of today, a &lt;a href=&#34;https://www.linkedin.com/search/results/people/?keywords=%22llm%20psychologist%22&#34;&gt;LinkedIn search for &amp;ldquo;llm psychologist&amp;rdquo;&lt;/a&gt; lists 9 people. I&amp;rsquo;m not alone!
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://www.linkedin.com/in/sanand0/&#34;&gt;Anand S&lt;/a&gt;, LLM Psychologist, Singapore, Singapore&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.linkedin.com/in/analyticsanshul/&#34;&gt;Anshul Saxena, PhD&lt;/a&gt;, AI Advisor &amp;amp; Trainer | Technology Strategist | LLM Psychologist | Currently teaching humans, machines &amp;amp; business to work smarter through Generative AI and Quantum Computing | 15+ Years Experience, Pune, Maharashtra, India&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.linkedin.com/in/chadofficial/&#34;&gt;Charitarth (Chad) Sindhu&lt;/a&gt;, LLM Psychologist / Fractional Business &amp;amp; AI Workflow Consultant/ Digital Nomad, Tokyo, Japan&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.linkedin.com/in/lancelotsalavert/&#34;&gt;Lancelot Salavert&lt;/a&gt;, LLM Psychologist, Barcelona, Catalonia, Spain&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.linkedin.com/in/lior-durahly/&#34;&gt;Lior Dor(Durahly)&lt;/a&gt;, Team Lead | Bug Banisher | Ex 8200, Tel Aviv District, Israel. Past: R&amp;amp;D Team Lead and &lt;strong&gt;LLM&lt;/strong&gt; &lt;strong&gt;Psychologist&lt;/strong&gt; at Superwise | A Blattner Tech Company&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.linkedin.com/in/maximebodereau/&#34;&gt;maxime bodereau&lt;/a&gt;, Lead Creative Art Director | UX Forensics | Ai LLM Psychologist | Visual Alchemist | Codesmith | Brandologist | Full Stack Designer, Nantes, Pays de la Loire, France&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.linkedin.com/in/chenml/&#34;&gt;Mei Chen 🦋&lt;/a&gt;, LLM Psychologist | Lead Product Engineer | Delivering Agentic Experiences, Toronto, Ontario, Canada&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.linkedin.com/in/shoshannahtekofsky/&#34;&gt;Shoshannah Tekofsky&lt;/a&gt;, LLM Psychologist at AI Digest, Zwolle, Overijssel, Netherlands&lt;/li&gt;
&lt;li&gt;LinkedIn Member, LLM, psychologist, mediator, Prague, Czechia&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://simonwillison.net/2026/Mar/19/openai-acquiring-astral/&#34;&gt;OpenAI acquired Astral!&lt;/a&gt;. This will likely slow down the new wonderful tools accelerating the Python ecosystem. Like with &lt;a href=&#34;https://openai.com/index/openai-to-acquire-promptfoo/&#34;&gt;PromptFoo&lt;/a&gt; and &lt;a href=&#34;https://steipete.me/posts/2026/openclaw&#34;&gt;OpenClaw&lt;/a&gt;, this seems to be about talent. The &amp;ldquo;acqui-hire&amp;rdquo; mode seems a &lt;em&gt;clear&lt;/em&gt; niche career path now, and an alternative to getting hired (you get a much higher salary) or getting acquired (you take on much higher risk).&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.npmjs.com/package/quickjs-emscripten&#34;&gt;quickjs-emscripten&lt;/a&gt; lets you run isolated JS code securely in the browser, CloudFlare workers, NodeJS, and Deno. It compiles to WASM. @sebastianwessel/quickjs is a higher-level TS wrapper. &lt;a href=&#34;https://github.com/simonw/research/tree/main/javascript-sandboxing-research&#34;&gt;Simon Willison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://bramcohen.com/p/manyana&#34;&gt;Manyana&lt;/a&gt; is a CRDT based version control system. It sounds like a good idea but I&amp;rsquo;m sceptical because merge conflicts are a &amp;ldquo;what should I do&amp;rdquo; problem more than &amp;ldquo;how&amp;rdquo;. With &lt;a href=&#34;https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/&#34;&gt;agents doing more merge conflict management&lt;/a&gt;, I am not sure this will offer a concrete benefit - but probably no harm either.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://posttrainbench.thoughtfullab.com/&#34;&gt;LLMs are able post-train LLMs on new topics&lt;/a&gt;. They&amp;rsquo;re improving fast. &lt;a href=&#34;https://jack-clark.net/2026/03/16/importai-449-llms-training-other-llms-72b-distributed-training-run-computer-vision-is-harder-than-generative-text/&#34;&gt;Jack Clark&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.linkedin.com/search/results/people/?keywords=vibe+coding+fixer&#34;&gt;Vibe Coding Fixer&lt;/a&gt; and &lt;a href=&#34;https://www.linkedin.com/search/results/people/?keywords=ai+slop+cleaner&#34;&gt;AI Slop Cleaner&lt;/a&gt; are real job descriptions - which are morphing into enterprise offerings. But I still seem to be the only official &lt;a href=&#34;https://www.linkedin.com/search/results/people/?keywords=llm+psychologist&#34;&gt;LLM Psychologist&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Notes from &lt;a href=&#34;https://mtrajan.substack.com/p/ai-services-wrong-mental-models-right&#34;&gt;AI Services - Wrong Mental Models, Right Moment&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;AI services has 3 markets. Automatable work: vanishes in 2 years. Human-in-the-loop work: sustains. Judgement-driven: grows in importance.&lt;/li&gt;
&lt;li&gt;YC: don’t sell access to a tool for $50 a month, use the AI yourself and sell the finished work for $5,000.&lt;/li&gt;
&lt;li&gt;Sell output. Price on outcome. Sell to business, not IT.&lt;/li&gt;
&lt;li&gt;Sell accountability: proven success, with your guarantee.&lt;/li&gt;
&lt;li&gt;Sell authenticity: a brand story representing uniqueness, character, &amp;hellip; or whatever&amp;hellip; something people respect.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Data transfer between GPU and memory is a bottleck and three approaches are emerging. &lt;a href=&#34;https://mtrajan.substack.com/p/inference-blindness&#34;&gt;#&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://taalas.com/the-path-to-ubiquitous-ai/&#34;&gt;Taalas&lt;/a&gt; is etching LLMs into the chip. Llama 8b runs at 17,000 tok/s (H200 is at 230 tok/s).&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.d-matrix.ai/announcements/d-matrix-unveils-corsair-the-worlds-most-efficient-ai-computing-platform-for-inference-in-datacenters/&#34;&gt;d-Matrix&lt;/a&gt; is moving compute into SRAM memory chips. 30,000 tok/s for Llama 70b. Cerebras and MatX are similar: memory-oriented.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://furiosa.ai/blog/lg-ai-research-taps-furiosaai-to-achieve-2-25x-better-llm-inference-in-production-vs-gpus&#34;&gt;FuriosaAI&lt;/a&gt; minimizes data movement. Groq and Sambanova are similar.&lt;/li&gt;
&lt;li&gt;But in the long run, commodity technology usually beats integrated stacks.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://openai.com/index/introducing-gpt-5-4-mini-and-nano/&#34;&gt;GPT 5.4 Nano ($0.2/MTok) and Mini ($0.75/MTok)&lt;/a&gt; are good options for bulk OCR, transcription, etc. as cost and quality comparable alternatives to Gemini Flash Lite and Gemini Flash. &lt;a href=&#34;https://simonwillison.net/2026/Mar/17/mini-and-nano/&#34;&gt;They can describe 75K photos for $50&lt;/a&gt;. Both models are better than GPT-5 Mini on most benchmarks.&lt;/li&gt;
&lt;li&gt;Cool &lt;a href=&#34;https://simonwillison.net/guides/agentic-engineering-patterns/using-git-with-coding-agents/&#34;&gt;AI coding agent git prompt fragments&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;Use git bisect to find when this bug was introduced: &amp;hellip;&lt;/li&gt;
&lt;li&gt;Find and recover my code that does &amp;hellip;&lt;/li&gt;
&lt;li&gt;Sort out this git mess for me.&lt;/li&gt;
&lt;li&gt;Rewrite history removing &amp;hellip;&lt;/li&gt;
&lt;li&gt;Split the last commit into multiple commits grouped logically.&lt;/li&gt;
&lt;li&gt;Start a new repo at &amp;hellip; and build just this module &amp;hellip; based on &amp;hellip; with a similar commit history copying the author and commit dates.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://matthodges.com/posts/2026-01-07-ai-agents-campaigns/&#34;&gt;Campaigns Are Knowledge Workers and the Tools Just Caught Up&lt;/a&gt;. A powerful framing. I saw this in action a few days ago when a friend was able to automate an outbound campaign with Claude Code.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.google.com/search?q=EARS+(Easy+Approach+to+Requirements+Syntax)&amp;amp;oq=EARS+(Easy+Approach+to+Requirements+Syntax)&#34;&gt;EARS (Easy Approach to Requirements Syntax)&lt;/a&gt; is a simple structure for requirements. For &lt;a href=&#34;https://github.com/github/spec-kit/issues/1356&#34;&gt;example&lt;/a&gt;, &amp;ldquo;Users should be able to drag tasks between columns. The app needs to work offline too. Handle errors gracefully.&amp;rdquo; becomes the following - which AI can convert to and is easier to spot errors in. State machines and decision tables are useful alternatives, too.
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;REQ-001&lt;/strong&gt; (Event): When the user drags a task card to a different column, the system shall update the task status to match the destination column.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;REQ-002&lt;/strong&gt; (State): While the application is offline, the system shall store task updates in local storage.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;REQ-003&lt;/strong&gt; (Event): When the application reconnects, the system shall synchronize locally stored updates with the server.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;REQ-004&lt;/strong&gt; (Unwanted): If synchronization conflicts occur, then the system shall display a resolution dialog to the user.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;As of now, avoid using Claude.ai to create (large) visualizations. It runs forever and exhausts credits without generating anything. Claude Code works much better for this.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>The Nov 2025 Vibe Coding Ghost Revolution</title>
      <link>https://www.s-anand.net/blog/the-nov-2025-vibe-coding-ghost-revolution/</link>
      <pubDate>Mon, 23 Mar 2026 11:21:42 +0530</pubDate>
      <guid>https://www.s-anand.net/blog/the-nov-2025-vibe-coding-ghost-revolution/</guid>
      <description>&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://sanand0.github.io/datastories/github-usage-increase/sketchnote.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;I kept hearing that with the &lt;a href=&#34;https://simonwillison.net/2026/Jan/4/inflection/&#34;&gt;Nov 2025 release&lt;/a&gt; of Opus 4.5 and GPT 5.2 Codex, &lt;a href=&#34;https://www.entrepreneur.com/business-news/this-ai-tool-is-helping-disempowered-ceos-with-a-major-problem-finally-feel-unleashed&#34;&gt;ex-coders&lt;/a&gt; &lt;a href=&#34;https://aimagazine.com/news/google-and-klarnas-ceos-are-vibe-coding-should-you-be&#34;&gt;were&lt;/a&gt; &lt;a href=&#34;https://technologymagazine.com/news/google-ceo-why-vibe-coding-makes-software-exciting-again&#34;&gt;sprinting&lt;/a&gt; &lt;a href=&#34;https://www.jpmorgan.com/insights/technology/artificial-intelligence/vibe-coding-a-guide-for-startups-and-founders&#34;&gt;back to coding&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;On a sample of ~1,700 developers on GitHub, exactly &lt;em&gt;ten&lt;/em&gt; fit the &amp;ldquo;dormant returner&amp;rdquo; profile.&lt;/p&gt;
&lt;p&gt;Here are a couple of examples:&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://github.com/tlwolsten&#34;&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-23-github-tlwolsten.webp&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://github.com/rjwalters&#34;&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-23-github-rjwalters.webp&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;But they&amp;rsquo;re the exception. I could find only &lt;strong&gt;TEN&lt;/strong&gt; out of 1,700 developers who returned. I also found a few who exited:&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://github.com/bocaletto-luca&#34;&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-23-github-bocaletto-luca.webp&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;To be fair, the vibe coding revolution &lt;em&gt;is&lt;/em&gt; real, but maybe we are (I am) mis-interpreting it.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;There are lots of new non-developers joining GitHub. Anecdotally, &lt;a href=&#34;https://github.com/gattun-git&#34;&gt;Naveen Gattu&lt;/a&gt; (finally!!) and Ankor Rai&lt;/li&gt;
&lt;li&gt;A few high-profile ex-developers are returning and are very active. Anecdotally, &lt;a href=&#34;https://en.wikipedia.org/wiki/Sebastian_Siemiatkowski&#34;&gt;Sebastian Siemiatkowski&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;But the majority of the developers who were less active last year remain less active.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;a href=&#34;https://sanand0.github.io/datastories/github-usage-increase/&#34;&gt;Read the full analysis&lt;/a&gt;.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Live Vibe Coding using Others&#39; Ideas</title>
      <link>https://www.s-anand.net/blog/live-vibe-coding-using-others-ideas/</link>
      <pubDate>Sat, 21 Mar 2026 22:51:50 +0530</pubDate>
      <guid>https://www.s-anand.net/blog/live-vibe-coding-using-others-ideas/</guid>
      <description>&lt;p&gt;I spoke today on &lt;a href=&#34;https://sanand0.github.io/talks/2026-03-21-design-in-the-age-of-infinite-generativity/&#34;&gt;Design in the Age of Infinite Generativity&lt;/a&gt; at the &lt;a href=&#34;https://www.chennaidesignfestival.com/&#34;&gt;Chennai Design Festival&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;You can read about the talk in the link about. This post is about my preparation.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-21-live-vibe-coding-using-others-ideas.avif&#34;&gt; &lt;!-- https://gemini.google.com/u/2/app/77c3c7f054bf447f --&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Tue 10 Mar 2026. Damn! &lt;a href=&#34;https://www.linkedin.com/in/palramu/&#34;&gt;Palani&lt;/a&gt;&amp;rsquo;s asked for the topic. &lt;a href=&#34;https://claude.ai/share/46ef75be-6e38-4ea6-ab73-49f7e46e1ec0&#34;&gt;Claude, what should I talk about!?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Fri 20 Mar 2026. &lt;a href=&#34;https://chatgpt.com/share/69bed4d0-14ac-8003-8750-2648b81d9366&#34;&gt;ChatGPT, tell me who the other speaker are&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Fri 20 Mar 2026. Oh, I&amp;rsquo;ll just pull a bunch of links, &lt;a href=&#34;https://www.s-anand.net/blog/using-browser-tabs-as-slides/&#34;&gt;use browser tabs as slides&lt;/a&gt;, create some &lt;a href=&#34;https://tools.s-anand.net/slide/&#34;&gt;slide dividers&lt;/a&gt;, and I&amp;rsquo;m ready!&lt;/li&gt;
&lt;li&gt;Sat 21 Mar 2026 1:00 pm. I&amp;rsquo;m &lt;strong&gt;NOT&lt;/strong&gt; ready! The story doesn&amp;rsquo;t flow. It&amp;rsquo;s rubbish.&lt;/li&gt;
&lt;li&gt;Sat 21 Mar 2026 3:00 pm. Let me drop some of the boring ones. I just have 15 minutes.&lt;/li&gt;
&lt;li&gt;Sat 22 Mar 2026 3:30 pm. Oh, maybe I should listen to what the others are saying, just&amp;hellip; you know&amp;hellip;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&amp;hellip; and that proved the &lt;em&gt;best&lt;/em&gt; decision ever, because &lt;a href=&#34;https://www.linkedin.com/in/senthil-gopalan-05935719/&#34;&gt;Senthil&lt;/a&gt; of &lt;a href=&#34;https://www.payir.org/&#34;&gt;Payir&lt;/a&gt; showed a &lt;a href=&#34;https://thinaistore.myinstamojo.com/product/fabric-calendar-hanging-model-reusable&#34;&gt;re-usable fabric calendar&lt;/a&gt; that converts into a bag. It was a fantastic idea, so I got curious.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;3:40 pm. &lt;a href=&#34;https://claude.ai/share/f9f8bce9-bb29-4795-b5c1-8faa7eca4dfe&#34;&gt;Asked Claude for more ideas like Senthil&amp;rsquo;s&lt;/a&gt;. It looked fine and I could read it later, but what if&amp;hellip; maybe&amp;hellip; I &lt;strong&gt;presented these ideas&lt;/strong&gt;!?&lt;/li&gt;
&lt;li&gt;3:45 pm. &lt;a href=&#34;https://gemini.google.com/share/c201f58b95b3&#34;&gt;Ask Gemini to draw the first idea - a Modular Kolam Mat&lt;/a&gt;. The results look &lt;em&gt;fantastic&lt;/em&gt;!
&lt;img loading=&#34;lazy&#34; src=&#34;https://sanand0.github.io/talks/2026-03-21-design-in-the-age-of-infinite-generativity/modular-kolam-1.avif&#34;&gt;&lt;/li&gt;
&lt;li&gt;3:50 pm. Now I&amp;rsquo;m going ga-ga over the idea. I generate images for 3 more ideas: a growth chart kurta, seed library sari border, and recipe towel.&lt;/li&gt;
&lt;li&gt;4:15 pm. &lt;a href=&#34;https://www.linkedin.com/in/narendraghate/&#34;&gt;Narendra&lt;/a&gt; shares a bunch of cool PsychOps design hacks like:
&lt;ul&gt;
&lt;li&gt;When lights are dimmed people speak softer. So, dimming lights reduces sound levels in noisy offices.&lt;/li&gt;
&lt;li&gt;Rather than reduce the size of shampoo sachets (which customers and business both hate), include 2 shampoos in one sachet, tearable in the middle.&lt;/li&gt;
&lt;li&gt;Price saches at 95p with a 5p deposit for the sachet - which rag-pickers can collect and return to the retailer.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;4:20 pm. &lt;a href=&#34;https://claude.ai/share/74c27359-7710-463e-8488-60fef2df6bfc&#34;&gt;Ask Claude for more ideas like Narendra&amp;rsquo;s&lt;/a&gt;. The results are &lt;em&gt;just as fantastic&lt;/em&gt;!
&lt;img loading=&#34;lazy&#34; src=&#34;https://sanand0.github.io/talks/2026-03-21-design-in-the-age-of-infinite-generativity/expired-medication-color.avif&#34;&gt;&lt;/li&gt;
&lt;li&gt;4:30 pm. I now have images for his ideas too. Now, I start deleting my more boring links.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In fact, these ideas ended up being &lt;em&gt;so&lt;/em&gt; good that the bulk of my talk was just about ideas derived from their work. (I&amp;rsquo;m obviously a big fan of plagiarism!)&lt;/p&gt;
&lt;p&gt;Despite being one of my shortest talks (~10 min), this worked quite well because people value speed and spontaneity. I was &lt;em&gt;obviously&lt;/em&gt; demonstrating these and that felt cool.&lt;/p&gt;
&lt;p&gt;For many years, I&amp;rsquo;ve been live-coding on stage. But that requires a &lt;em&gt;lot&lt;/em&gt; of preparation.&lt;/p&gt;
&lt;p&gt;Vibe-coding makes live-coding a &lt;em&gt;lot&lt;/em&gt; faster. I can do it &lt;em&gt;during&lt;/em&gt; a client demo. I can do it &lt;em&gt;during&lt;/em&gt; a talk.&lt;/p&gt;
&lt;p&gt;So, I&amp;rsquo;m going to listen more to what others are saying (in meetings, conferences, etc.) and live-vibe-code from what they &lt;em&gt;just&lt;/em&gt; said. Great way to show-off while learning from others!&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;26 Mar 2026&lt;/strong&gt;: On a related note, &lt;a href=&#34;https://x.com/emollick/status/2036104905568452967&#34;&gt;AI has better product development ideas than humans&lt;/a&gt;. Read &lt;a href=&#34;https://arxiv.org/abs/2603.19087&#34;&gt;Serendipity by Design: Evaluating the Impact of Cross-domain Mappings on Human and LLM Creativity&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Protyping the prototypes</title>
      <link>https://www.s-anand.net/blog/prototyping-the-prototypes/</link>
      <pubDate>Tue, 10 Mar 2026 12:17:47 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/prototyping-the-prototypes/</guid>
      <description>&lt;p&gt;I added a narrative story to my &lt;a href=&#34;https://sanand0.github.io/llmpricing/&#34;&gt;LLM Pricing chart&lt;/a&gt;. That makes it easier for me &lt;em&gt;and&lt;/em&gt; others to tell the story of AI&amp;rsquo;s evolution in the last three years.&lt;/p&gt;
&lt;video controls autoplay loop muted playsinline preload=&#34;metadata&#34; width=&#34;1400&#34; height=&#34;800&#34; style=&#34;max-width: 100%; height: auto;&#34;&gt;
  &lt;source src=&#34;https://files.s-anand.net/images/2026-03-10-llmpricing-screencast-crf55-fps5.webm&#34; type=&#34;video/webm&#34;&gt;
  &lt;a href=&#34;https://files.s-anand.net/images/2026-03-10-llmpricing-screencast-crf55-fps5.webm&#34;&gt;Video&lt;/a&gt;
&lt;/video&gt;
&lt;p&gt;It was vibe-coded over two iterations.&lt;/p&gt;
&lt;p&gt;In &lt;a href=&#34;https://github.com/sanand0/llmpricing/blob/467474abd9ebdb3051ba016ddc95bfed7da556c6/prompt.md#scrolly-v1&#34;&gt;the first version&lt;/a&gt;, I prompted it to:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Add a scrollytelling narrative. So, when users first visit the page, they see roughly the same thing as now (but prettier). As they scroll down, the page should smoothly move to the earliest month, and then animate month by month on scroll, and explaining the key events and insights in terms of model quality and pricing. Use the data story skill to do this effectively, narrating like Malcolm Gladwell, with the visual style of The New York Times, using the education progression as a framework for measure of intelligence (read prompts.md for context). Store the narrative text in a separate JSON file and read from it. This should control the entire narrative, including what month to jump to next, what models to highlight, what insights to share, and so on.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;But there were two problems:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Conflicting instructions&lt;/strong&gt;. &amp;ldquo;&amp;hellip; with the visual style of The New York Times&amp;rdquo; conflicted with my current style.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Incomplete instructions&lt;/strong&gt;. I wanted to begin with the exploration, not the narrative. I wanted to explain the axes first. I wanted smaller cards.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Both of them were solved in the &lt;a href=&#34;https://github.com/sanand0/llmpricing/blob/467474abd9ebdb3051ba016ddc95bfed7da556c6/prompt.md#scrolly-v2&#34;&gt;second version&lt;/a&gt;, because this time, &lt;em&gt;I knew what I wanted&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;In other words, vibe coding isn&amp;rsquo;t just helping me prototype, it&amp;rsquo;s helping me prototype the prototype! Even if I don&amp;rsquo;t know what I want, I can just ask for something and build on it.&lt;/p&gt;
&lt;p&gt;This is a known benefit of lower costs. But, like the placebo effect and Hofstadter&amp;rsquo;s Law, I&amp;rsquo;m surprised by it even when I know it.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;An aside: Here&amp;rsquo;s the &lt;a href=&#34;https://github.com/sanand0/llmpricing/blob/467474abd9ebdb3051ba016ddc95bfed7da556c6/narrative.json&#34;&gt;narrative&lt;/a&gt; it crafted:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;How to read this chart&lt;/strong&gt;: The vertical axis is intelligence — mapped to the academic ladder on the right. Elo 1100 is high school freshman; Elo 1480 is tenured professor. The horizontal axis is cost: one million input tokens — roughly the entire King James Bible — priced from two cents to $75. &lt;strong&gt;The upper-left corner is the dream: brilliant and cheap.&lt;/strong&gt; This is the story of how the world got there.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;March 2023 — Year Zero&lt;/strong&gt;: March 2023. The entire AI landscape fits in a small cluster near the bottom of the chart. GPT-3.5 and Claude 1 are the state of the art — high-school-to-college-freshman intelligence that writes a fluent paragraph, then confidently invents a fact. Processing the King James Bible costs 50 cents to $8. These models can hold a conversation. &lt;strong&gt;They cannot, reliably, hold an argument.&lt;/strong&gt; &lt;a href=&#34;https://lmsys.org/blog/2023-05-03-arena/&#34;&gt;LMSYS Chatbot Arena launches (May 2023)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;November 2023 — The Leap&lt;/strong&gt;: November 6, 2023: GPT-4 Turbo. The chart jolts upward. Elo 1313 — college-junior level: coherent research papers, complex reasoning, output worth reading. &lt;strong&gt;Price: $10 per million tokens&lt;/strong&gt; — about $14 to process all seven Harry Potter novels. Expensive, but for the first time the intelligence felt worth it. Enterprises stopped asking whether AI could help. They started asking how much they were willing to pay. &lt;a href=&#34;https://openai.com/blog/new-models-and-developer-products-announced-at-devday&#34;&gt;OpenAI DevDay announcement&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;February 2024 — The Split&lt;/strong&gt;: Early 2024: Anthropic launches Claude 3 in three tiers. Opus and Sonnet arrive February 29; Haiku a week later, March 7. Opus: entry-analyst capability at $15. Sonnet: similar quality at $3. Haiku: college-junior reasoning at $0.25. &lt;strong&gt;A 60× price spread — same company, same training philosophy.&lt;/strong&gt; The intelligence market had learned to stratify, and every business began thinking in tiers. &lt;a href=&#34;https://www.anthropic.com/news/claude-3-family&#34;&gt;Anthropic Claude 3 announcement&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;June 2024 — The Summer Pivot&lt;/strong&gt;: June 2024 bent the economics permanently. Claude 3.5 Sonnet — strong manager level (Elo 1342) — arrived at $3 per million tokens: five times cheaper than Opus, four months later, at better quality. Meta’s Llama 3.1 405B matched it at $2, or free if self-hosted. Any CFO paying $15 for frontier AI could now pay $3. &lt;strong&gt;The question changed from “can we afford AI?” to “what are we waiting for?”&lt;/strong&gt; &lt;a href=&#34;https://www.anthropic.com/news/claude-3-5-sonnet&#34;&gt;Claude 3.5 Sonnet launch&lt;/a&gt;, &lt;a href=&#34;https://ai.meta.com/blog/meta-llama-3-1/&#34;&gt;Meta releases Llama 3.1&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;September 2024 — The Thinking Machine&lt;/strong&gt;: September 2024: OpenAI o1. It didn’t just answer — it reasoned. Before responding, it ran an internal monologue: checking its own logic, catching its own errors. &lt;strong&gt;Elo 1388 — the biggest single-model quality jump since GPT-4.&lt;/strong&gt; It scored at or above PhD-expert level on GPQA Diamond, a benchmark of graduate-level science questions. Price: $15. For the first time, AI felt less like autocomplete and more like a colleague you’d genuinely consult. &lt;a href=&#34;https://openai.com/index/openai-o1-system-card/&#34;&gt;OpenAI o1 system card&lt;/a&gt;, &lt;a href=&#34;https://arxiv.org/abs/2311.12022&#34;&gt;GPQA Diamond benchmark results&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;January 2025 — The Earthquake&lt;/strong&gt;: January 20, 2025. DeepSeek R1 matched o1-level reasoning (Elo 1398) at $0.55 per million tokens — &lt;strong&gt;a 27× discount.&lt;/strong&gt; The announcement wiped $600 billion from Nvidia’s market cap in a single day. Silicon Valley assumed expensive compute was a moat. DeepSeek proved it was just a starting point. PhD-approaching reasoning for fifty-five cents per Bible. &lt;a href=&#34;https://www.reuters.com/technology/chinas-deepseek-sets-off-ai-market-rout-2025-01-27/&#34;&gt;Reuters: DeepSeek wipes $600B from Nvidia&lt;/a&gt;, &lt;a href=&#34;https://arxiv.org/abs/2501.12948&#34;&gt;DeepSeek R1 technical report&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mid-2025 — The Race to the Top-Left&lt;/strong&gt;: By mid-2025, the upper-left corner was filling fast. Gemini 2.5 Pro delivered tenured-professor intelligence (Elo 1476) at $1.25 per million tokens. Flash models handled most enterprise work for 30 cents. Companies that had rationed AI to critical workflows were now running it everywhere. &lt;strong&gt;The constraint was no longer cost or capability — it was imagination.&lt;/strong&gt; &lt;a href=&#34;https://blog.google/technology/google-deepmind/gemini-model-updates-february-2025/&#34;&gt;Google Gemini 2.5 Pro launch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;February 2026 — Polymath Scholar&lt;/strong&gt;: The top models in early 2026 score above 1500 Elo — polymath scholar territory. A 1500-rated model beats a 1300-rated one in 76% of head-to-head matchups. GPT-4 Turbo launched at college-junior level (Elo 1313) in November 2023. Twenty-seven months later, Claude Opus 4.6 and Gemini 3.1 Pro sit at 1500+. &lt;strong&gt;What cost $15 in 2024 is now beaten by models at one-seventh the price.&lt;/strong&gt; &lt;a href=&#34;https://lmarena.ai/leaderboard&#34;&gt;LMArena Leaderboard&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;So What Does This Mean?&lt;/strong&gt;: In 2023, reaching master’s-level AI for 100 analysts meant paying $30/MTok — roughly $135,000 a month. By 2026, models with higher capability cost $1–5/MTok: under $15,000 for better results. For most tasks — analysis, drafting, coding, research — models in the $1–3 range deliver more than enough. Save frontier models for your hardest problems. &lt;strong&gt;The question is no longer what AI costs. It’s what you would build if intelligence cost a dime per Bible.&lt;/strong&gt; &lt;a href=&#34;https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential-of-generative-ai&#34;&gt;McKinsey: The economic potential of generative AI&lt;/a&gt;, &lt;a href=&#34;https://www.sequoiacap.com/article/ais-600b-question/&#34;&gt;Sequoia: AI&amp;rsquo;s $600B question&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;p&gt;Self-referentially, my answer to its last question: &amp;ldquo;what you would build if intelligence cost a dime per Bible&amp;rdquo; is &amp;ldquo;a scrollytelling narrative about intellegince costing a dime per Bible&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-10-prototyping-the-prototypes.avif&#34;&gt; &lt;!-- https://gemini.google.com/u/2/app/6b7cf5190d354c87 --&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>AWS PartyRock</title>
      <link>https://www.s-anand.net/blog/aws-partyrock/</link>
      <pubDate>Thu, 22 Jan 2026 21:22:31 +0530</pubDate>
      <guid>https://www.s-anand.net/blog/aws-partyrock/</guid>
      <description>&lt;p&gt;I tried vibe-code a CSV to colored HTML table converter using &lt;a href=&#34;https://github.com/sanand0/tools/blob/09ad622f1bfb662a5203593836ef765b99839e8a/colortable/prompts.md&#34;&gt;this prompt&lt;/a&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Create a tool that can convert pasted tables into colored HTML tables.&lt;/li&gt;
&lt;li&gt;Allow the user to paste a CSV or tab-delimited or pipe-delimited table.&lt;/li&gt;
&lt;li&gt;&amp;hellip;&lt;/li&gt;
&lt;li&gt;Create an HTML table that has minimal styling.&lt;/li&gt;
&lt;li&gt;&amp;hellip;&lt;/li&gt;
&lt;li&gt;Add a button to copy just the HTML to the clipboard.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href=&#34;https://tools.s-anand.net/colortable/&#34;&gt;Codex built this&lt;/a&gt;. Which is perfect.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://partyrock.aws/u/sanand0/Umcv1aE1W/TableMorph%253A-CSV-to-Styled-HTML-Converter&#34;&gt;AWS Partyrock built this&lt;/a&gt;. Which is a joke, because it didn&amp;rsquo;t write the code to do the conversion. It uses an LLM every time.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-01-22-aws-partyrock.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;The good part, though, is that it builds UI components that you can edit and move around.&lt;/p&gt;
&lt;p&gt;I think it builds a domain specific language with components with properties. You can edit the properties to change the app.&lt;/p&gt;
&lt;p&gt;This is a good way to build robust apps using unreliable LLMs. But it has less flexibility and richness.&lt;/p&gt;
&lt;p&gt;With LLMs becoming more and more reliable, I think generating code directly is better, and this approach will fade or be used in niches.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>NPTEL Applied Vibe Coding Workshop</title>
      <link>https://www.s-anand.net/blog/nptel-applied-vibe-coding-workshop/</link>
      <pubDate>Sun, 11 Jan 2026 22:53:28 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/nptel-applied-vibe-coding-workshop/</guid>
      <description>&lt;p&gt;For those who missed my &lt;a href=&#34;https://elearn.nptel.ac.in/shop/iit-workshops/ongoing/computer-science/applied-vibe-coding-workshop/&#34;&gt;Applied Vibe Coding Workshop&lt;/a&gt; at NPTEL, here&amp;rsquo;s the video:&lt;/p&gt;
&lt;div class=&#34;video-embed&#34;&gt;&lt;iframe src=&#34;https://www.youtube.com/embed/m9mIe4baN-k&#34; title=&#34;YouTube video&#34; loading=&#34;lazy&#34; allow=&#34;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture&#34; allowfullscreen&gt;&lt;/iframe&gt;&lt;/div&gt;
&lt;p&gt;You can also:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://sanand0.github.io/talks/2026-01-11-nptel-vibe-coding-workshop/&#34;&gt;Read this summary of the talk&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/talks/blob/main/2026-01-11-nptel-vibe-coding-workshop/transcript.md&#34;&gt;Read the transcript&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img alt=&#34;Sketchnote of the talk&#34; loading=&#34;lazy&#34; src=&#34;https://raw.githubusercontent.com/sanand0/talks/refs/heads/main/2026-01-11-nptel-vibe-coding-workshop/sketchnote.avif&#34;&gt;&lt;/p&gt;
&lt;p&gt;Or, here are the three dozen lessons from the workshop:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Definition: Vibe coding is building apps by talking to a computer instead of typing thousands of lines of code.&lt;/li&gt;
&lt;li&gt;Foundational Mindset Lessons
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;In a workshop, you do the work&amp;rdquo;&lt;/strong&gt; - Learning happens through doing, not watching.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;If I say something and AI says something, trust it, don&amp;rsquo;t trust me&amp;rdquo;&lt;/strong&gt; - For factual information, defer to AI over human intuition.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Don&amp;rsquo;t ever be stuck anywhere because you have something that can give you the answer to almost any question&amp;rdquo;&lt;/strong&gt; - AI eliminates traditional blockers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Imagination becomes the bottleneck&amp;rdquo;&lt;/strong&gt; - Execution is cheap; knowing &lt;em&gt;what&lt;/em&gt; to build is the constraint.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Doing becomes less important than knowing what to do&amp;rdquo;&lt;/strong&gt; - Strategic thinking outweighs tactical execution.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;You don&amp;rsquo;t have to settle for one option. You can have 20 options&amp;rdquo;&lt;/strong&gt; - AI makes parallel exploration cheap.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Practical Vibe Coding Lessons
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Success metric&lt;/strong&gt;: &amp;ldquo;Aim for 10 applications in a 1-2 hour workshop&amp;rdquo; - Volume and iteration over perfection.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The subscription vs. platform distinction&lt;/strong&gt;: &amp;ldquo;Your subscriptions provide the brains to write code, but don&amp;rsquo;t give you tools to host and turn it into a live working app instantly.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Add documentation for users&lt;/strong&gt;: First-time users need visual guides or onboarding flows.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Error fixing success rate&lt;/strong&gt;: &amp;ldquo;About one in three times&amp;rdquo; fixing errors works. &amp;ldquo;If it doesn&amp;rsquo;t work twice, start again-sometimes the same prompt in a different tab works.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Planning mode before complex builds&lt;/strong&gt;: &amp;ldquo;Do some research. Find out what kind of application along this theme can be really useful and why. Give me three or four options.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ask &amp;ldquo;Do I need an app, or can the chatbot do it?&amp;rdquo;&lt;/strong&gt; - Sometimes direct AI conversation beats building an app.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Local HTML files work&lt;/strong&gt;: &amp;ldquo;Just give me a single HTML file&amp;hellip; opening it in my browser should work&amp;rdquo; - No deployment infrastructure needed.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;The skill we are learning is how to learn&amp;rdquo;&lt;/strong&gt; - Specific tool knowledge is temporary; meta-learning is permanent.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Vibe Analysis Lessons
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;The most interesting data sets are our own data&amp;rdquo;&lt;/strong&gt; - Personal data beats sample datasets.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Accessible personal datasets&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;WhatsApp chat exports&lt;/li&gt;
&lt;li&gt;Netflix viewing history (Account &amp;gt; Viewing Activity &amp;gt; Download All)&lt;/li&gt;
&lt;li&gt;Local file inventory (&lt;code&gt;ls -R&lt;/code&gt; or equivalent)&lt;/li&gt;
&lt;li&gt;Bank/credit card statements&lt;/li&gt;
&lt;li&gt;Screen time data (screenshot &amp;gt; AI digitization)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ChatGPT&amp;rsquo;s hidden built-in tools&lt;/strong&gt;: FFmpeg (audio/video), ImageMagick (images), Poppler (PDFs)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Code as art form&amp;rdquo;&lt;/strong&gt; - Algorithmic art (Mandelbrot, fractals, Conway&amp;rsquo;s Game of Life) can be AI-generated and run automatically.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Data stories vs dashboards&amp;rdquo;&lt;/strong&gt;: &amp;ldquo;A dashboard is basically when we don&amp;rsquo;t know what we want.&amp;rdquo; Direct questions get better answers than open-ended visualization.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Prompting Wisdom
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Analysis prompt framework&lt;/strong&gt;: &amp;ldquo;Analyze data like an investigative journalist&amp;rdquo; - find surprising insights that make people say &amp;ldquo;Wait, really?&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cross-check prompt&lt;/strong&gt;: &amp;ldquo;Check with real world. Check if you&amp;rsquo;ve made a mistake. Check for bias. Check for common mistakes humans make.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Visualization prompt&lt;/strong&gt;: &amp;ldquo;Write as a narrative-driven data story. Write like Malcolm Gladwell. Draw like the New York Times data visualization team.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;20 years of experience&amp;rdquo;&lt;/strong&gt; - Effective prompts require domain expertise condensed into instructions.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Security &amp;amp; Governance
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Simon Willison&amp;rsquo;s &amp;ldquo;Lethal Trifecta&amp;rdquo;&lt;/strong&gt;: Private data + External communication + Untrusted content = Security risk. &lt;strong&gt;Pick any two, never all three.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;What constitutes untrusted content is very broad&amp;rdquo;&lt;/strong&gt; - Downloaded PDFs, copy-pasted content, even AI-generated text may contain hidden instructions.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Same governance as human code&lt;/strong&gt;: &amp;ldquo;If you know what a lead developer would do to check junior developer code, do that.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Treat AI like an intern&lt;/strong&gt;: &amp;ldquo;The way I treat AI is exactly the way I treat an intern or junior developer.&amp;rdquo;&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Business &amp;amp; Career Implications
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Social skills have a higher uplift on salary than math or engineering skills&amp;rdquo;&lt;/strong&gt; - Research finding from mid-80s/90s onward.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Differentiation challenge&lt;/strong&gt;: &amp;ldquo;If you can vibe code, anyone can vibe code. The differentiation will come from the stuff you are NOT vibe coding.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;The highest ROI investment I&amp;rsquo;ve made in life is paying $20 for ChatGPT or Claude&amp;rdquo;&lt;/strong&gt; - Worth more than 30 Netflix subscriptions in utility.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Where Vibe Coding Fails
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Failure axes&lt;/strong&gt;: &amp;ldquo;Large&amp;rdquo; and &amp;ldquo;not easy for software to do&amp;rdquo; - Complexity increases failure rates.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Local LLMs (Ollama, etc.)&lt;/strong&gt;: &amp;ldquo;Possible but not as fast or capable. Useful offline, but doesn&amp;rsquo;t match online experience yet.&amp;rdquo;&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Final Takeaways
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Practice vibe coding every day for one month&amp;rdquo;&lt;/strong&gt; - Habit formation requires forced daily practice.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Learn to give up&amp;rdquo;&lt;/strong&gt; - When something fails repeatedly, start fresh rather than debugging endlessly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Share what you vibe coded&amp;rdquo;&lt;/strong&gt; - Teaching others cements your own learning. &amp;ldquo;We learn best when we teach.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool knowledge is temporary&lt;/strong&gt;: &amp;ldquo;This field moves so fast, by the time somebody comes up with a MOOC, it&amp;rsquo;s outdated.&amp;rdquo;&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title></title>
      <link>https://www.s-anand.net/blog/nano-banana-pro-text-generation/</link>
      <pubDate>Sat, 22 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/nano-banana-pro-text-generation/</guid>
      <description>&lt;p&gt;Nano Banano Pro has &lt;strong&gt;excellent&lt;/strong&gt; text generation (though it doesn&amp;rsquo;t always give you what you want in the first try).&lt;/p&gt;
&lt;p&gt;I couldn&amp;rsquo;t spot any errors in the generated text. Can you?&lt;/p&gt;
&lt;p&gt;I used this prompt (with the workshop details and my photo): &lt;em&gt;Create a professional poster for the below, including all relevant information. Use my photo (attached) professionally&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;The NPTEL workshop is real, BTW. First 100 seats, I think. You can register here: &lt;a href=&#34;https://elearn.nptel.ac.in/shop/iit-workshops/ongoing/computer-science/applied-vibe-coding-workshop/&#34;&gt;https://elearn.nptel.ac.in/shop/iit-workshops/ongoing/computer-science/applied-vibe-coding-workshop/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2025-11-22-nptel-applied-vibe-coding-workshop-linkedin.jpg&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.linkedin.com/posts/sanand0_nano-banano-pro-has-%F0%9D%97%B2%F0%9D%98%85%F0%9D%97%B0%F0%9D%97%B2%F0%9D%97%B9%F0%9D%97%B9%F0%9D%97%B2%F0%9D%97%BB%F0%9D%98%81-text-activity-7398533321317769216-hezE&#34;&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 26 Oct 2025</title>
      <link>https://www.s-anand.net/blog/things-i-learned-26-oct-2025/</link>
      <pubDate>Sun, 26 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-26-oct-2025/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Before founding a place to do good, work in a place that does good and learn. &lt;a href=&#34;https://werd.io/using-technology-skills-for-positive-change/&#34;&gt;Ben Werdmuller&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;What should we teach when vibe coding becomes good enough for non-coders? &lt;a href=&#34;https://x.com/emollick/status/1979627762903392362&#34;&gt;Ethan Mollick&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Problem decomposition&lt;/li&gt;
&lt;li&gt;Clear communication &amp;amp; spec writing&lt;/li&gt;
&lt;li&gt;Core technical foundations: file systems, access control, networking, APIs, version control, data structures, databases, deployment&lt;/li&gt;
&lt;li&gt;Software development skills: Debugging, Testing, Refactoring, Design patterns, UI/UX&lt;/li&gt;
&lt;li&gt;Project management: requirements, prioritization, scoping, &amp;hellip;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Codex CLI tips:
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;codex --add-dir $DIR&lt;/code&gt; lets you write into $DIR&lt;/li&gt;
&lt;li&gt;&lt;code&gt;codex --full-auto&lt;/code&gt; is the equivalent of &lt;code&gt;codex --sandbox workspace-write --ask-for-approval on-request&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Terse code is not necessarily easier or harder for LLMs to write. It&amp;rsquo;s about how unusual (or not aligned with training data) the code is. &lt;a href=&#34;https://medium.com/@gabiteodoru/dont-force-your-llm-to-write-terse-code-an-argument-from-information-theory-for-q-kdb-developers-04077c5b7038&#34;&gt;Gabi Teoduru&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;How are people using browser agents like Comet / Atlas? &lt;a href=&#34;https://x.com/simonw/status/1980713097024401548&#34;&gt;Simon Willison&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Most popular: YouTube video summaries with timestamps&lt;/li&gt;
&lt;li&gt;Most useful: Form filling: Government forms, data entry, repetitive bureaucratic tasks
&lt;ul&gt;
&lt;li&gt;Foreign language navigation: Applying for pension in Korea, navigating sites in other languages&lt;/li&gt;
&lt;li&gt;Time reporting auto-completion&lt;/li&gt;
&lt;li&gt;Insurance claims: Reading policy documents and drafting appeals (successfully got claim reimbursed in India)&lt;/li&gt;
&lt;li&gt;Compliance training click throughs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Next most useful: Shopping / planning
&lt;ul&gt;
&lt;li&gt;Energy provider comparison - Comet checked current plan vs competitors on Check24, calculated exact annual savings per provider&lt;/li&gt;
&lt;li&gt;Financial tracking: Finding Amazon orders, tracking Airbnb spending with refund calculations, analyzing bank transactions&lt;/li&gt;
&lt;li&gt;Trip planning: Mapping 50-100 places on Google Maps automatically&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Interesting: Airport shuttle discovery - Found shuttle that user missed in manual searching&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/winfsp/hubfs&#34;&gt;HubFS&lt;/a&gt; mounts GitHub repos on the file system. Every file system action directly works on GitHub via a REST API. Useful for some scenarios but less useful for note-taking than something like &lt;a href=&#34;https://marketplace.visualstudio.com/items?itemName=vsls-contrib.gitdoc&#34;&gt;GitDoc&lt;/a&gt; which offers a delayed sync.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://x.com/ErnestRyu/status/1980760351479328781&#34;&gt;Ernest Ryu solved an open problem in convex optimization using ChatGPT&lt;/a&gt;. Quotes:
&lt;ul&gt;
&lt;li&gt;ChatGPT is now at the level of solving some math research questions, but you do need an expert guiding it.&lt;/li&gt;
&lt;li&gt;ChatGPT was really effective at accelerating my progress. This work took about 12 hours, spread over 3 days. In hindsight, the proof is really simple.&lt;/li&gt;
&lt;li&gt;But I iterated through so many other strategies that didn&amp;rsquo;t pan out, and ChatGPT crucially helped to quickly explore and eliminate those dead-end approaches. Also, the key successful steps were suggested by ChatGPT.&lt;/li&gt;
&lt;li&gt;ChatGPT did not produce the proof in a single prompt. The process was highly interactive. It generated many arguments, roughly 80% of which were incorrect.&lt;/li&gt;
&lt;li&gt;Yet some were genuinely novel to me. Whenever I recognized a novel idea, whether correct or only partially so, I distilled the key insight and prompted ChatGPT to develop it further.&lt;/li&gt;
&lt;li&gt;My contribution:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Filtering out incorrect arguments&lt;/strong&gt; and accumulating a set of correct facts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Identifying promising new lines&lt;/strong&gt; of reasoning and guiding ChatGPT to explore them further&lt;/li&gt;
&lt;li&gt;Recognizing when a strategy had been fully explored and &lt;strong&gt;deciding when to move on&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;ChatGPT&amp;rsquo;s contribution:
&lt;ul&gt;
&lt;li&gt;Producing the final proof argument.&lt;/li&gt;
&lt;li&gt;Significantly accelerating my (or our) exploration of the many dead-end arguments, rapidly ruling out approaches that did not work.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Comparing the GPT 4.1 and 5 models at all different of reasoning, I&amp;rsquo;ve switched my default from GPT 4.1 mini to GPT 5 mini (medium). Far smarter for a slightly higher cost. &lt;a href=&#34;https://artificialanalysis.ai/?cost=cost-vs-intelligence&amp;amp;models=gpt-5-low%2Cgpt-5-minimal%2Cgpt-5-nano%2Cgpt-5-nano-minimal%2Cgpt-5-mini%2Cgpt-5%2Cgpt-5-medium%2Cgpt-5-nano-medium%2Cgpt-5-mini-minimal%2Cgpt-5-mini-medium%2Capriel-v1-5-15b-thinker%2Cgpt-4-1%2Cgpt-4-1-nano%2Cgpt-4-1-mini&#34;&gt;Artificial Analysis&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;python -m pdb -c continue script.py&lt;/code&gt; or &lt;code&gt;uv run -m pdb -c continue script.py&lt;/code&gt; runs a script and drops into pdb on unhandled exceptions (post-mortem). &lt;a href=&#34;https://chatgpt.com/share/68f9b890-ba0c-800c-8a29-48245a41ca5e&#34;&gt;ChatGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Technology removes constraints. We then do what we really value. &lt;a href=&#34;https://claude.ai/chat/f3a2606f-203c-41cc-b50f-62504483504f&#34;&gt;Claude&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;When writing became digitized, we stopped cared about spelling/handwriting for its own sake. Spelling bees and handwriting classes declined. &amp;ldquo;ur&amp;rdquo; is acceptable.&lt;/li&gt;
&lt;li&gt;When fitness tracking became easy, many just track, few exercise more. Few people value exercise&lt;/li&gt;
&lt;li&gt;When GPS became ubiquitous, we stopped learning geography. Most value arriving, not knowing&lt;/li&gt;
&lt;li&gt;When photography became unlimited, most captured moments. Few perfected shots&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;I had Codex scrape my ~2,000 pending invites on LinkedIn and asked ChatGPT to analyze it. Here are learnings: &lt;a href=&#34;https://chatgpt.com/c/68f72899-5814-8320-9d02-88ce06257fd8&#34;&gt;ChatGPT, private&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Power-law. 5% of inviters account for ~42% of all common connections. Top 10 people alone for ~20%.&lt;/li&gt;
&lt;li&gt;IITM student invites are high (~14%), but with 0-2 common connects, i.e. distant strangers.&lt;/li&gt;
&lt;li&gt;EdTech is tiny in count but has the highest common connections per person (outlier-sensitive but real).&lt;/li&gt;
&lt;li&gt;Among ≥20-commons, many hold VP/Head/Site-Lead titles in Data/AI or GenAI (not just recruiters).&lt;/li&gt;
&lt;li&gt;GenAI people are 7-8% and steady across months. Not a useful signal to prioritize.&lt;/li&gt;
&lt;li&gt;Premium ~ Senior. Premium accounts show ~40% senior titles vs ~29% for non-premium.&lt;/li&gt;
&lt;li&gt;Finance invites have higher seniority rate and more common connects than healthcare.&lt;/li&gt;
&lt;li&gt;Followers have higher common connections (~6 vs ~4).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;⭐ Memory can be code. Agent memory is anything it choose to persist. Agents can write code on the fly to automate tasks, save them, and serve the code on the next request, potentially modifying the code as required. This is like the conscious mind saving a habit for the subconscious to execute fast.&lt;/li&gt;
&lt;li&gt;Finally: Microsoft Office has an agent mode that lets you talk to it and do stuff. &lt;a href=&#34;https://www.theverge.com/news/787076/microsoft-office-agent-mode-office-agent-anthropic-models&#34;&gt;The Verge&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Vibe-Coding for Interesting Data Stories</title>
      <link>https://www.s-anand.net/blog/vibe-coding-for-interesting-data-stories/</link>
      <pubDate>Mon, 06 Oct 2025 09:03:35 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/vibe-coding-for-interesting-data-stories/</guid>
      <description>&lt;p&gt;&lt;img alt=&#34;Vibe-Coding for Interesting Data Stories&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/gardener.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;Last weekend, I fed Codex my browser history and said &amp;ldquo;explore.&amp;rdquo; It found a pattern I call &lt;strong&gt;rabbit holes&lt;/strong&gt; &amp;ndash; three ways we browse:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Linear spiral&lt;/strong&gt; - one page &amp;gt; next page &amp;gt; next. E.g. filing income tax, clicking &amp;ldquo;next&amp;rdquo; on the &lt;a href=&#34;https://in.pycon.org/2025/program/schedule/&#34;&gt;PyCon schedule&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hub &amp;amp; spoke&lt;/strong&gt; - hub &amp;gt; open tabs &amp;gt; back to hub. E.g. exploring &lt;a href=&#34;https://en.wikipedia.org/wiki/David_Heinemeier_Hansson&#34;&gt;DHH&lt;/a&gt;&amp;rsquo;s Ubuntu setup, checking Firebase config.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Wide survey&lt;/strong&gt; - source &amp;gt; many, many pages. E.g. clearing inbox, scanning news.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Then Claude Code built this &lt;a href=&#34;https://sanand0.github.io/datastories/browser-history/rabbit-holes/&#34;&gt;lovely data story&lt;/a&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;My goal? Find challenges in vibe-coding &lt;strong&gt;interesting&lt;/strong&gt; data stories. I found several.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A. I don&amp;rsquo;t know what I want.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Solution? &lt;strong&gt;Ask for multiple options&lt;/strong&gt;. More options = more ideas. Codex proposed two I hadn&amp;rsquo;t planned: &lt;a href=&#34;https://sanand0.github.io/datastories/browser-history/rabbit-holes/&#34;&gt;rabbit holes&lt;/a&gt; and &lt;a href=&#34;https://sanand0.github.io/datastories/browser-history/search-funnels/&#34;&gt;search funnels&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;B. I don&amp;rsquo;t know if it&amp;rsquo;ll turn out well.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Solution? &lt;strong&gt;Build them all&lt;/strong&gt;. Don&amp;rsquo;t pre-judge. I &lt;strong&gt;did not&lt;/strong&gt; expect rabbit holes to be interesting - a clear prediction error.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;C. Reviewing is the bottleneck.&lt;/strong&gt; It&amp;rsquo;s slow and painful.&lt;/p&gt;
&lt;p&gt;Solution? &lt;strong&gt;Make reviews easy&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Ask for review-friendly output&lt;/strong&gt;. E.g. A table/heatmap comparing options.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Use LLMs to pre-review&lt;/strong&gt;. E.g. Pick top 3 with reasons.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Review output, not code&lt;/strong&gt;. E.g. Have it build a working demo, &lt;strong&gt;then&lt;/strong&gt; review.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;D. Model / tool strengths vary.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Solution? &lt;strong&gt;Align with strengths&lt;/strong&gt;. For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Use GPT-5 for planning&lt;/strong&gt;. It&amp;rsquo;s better than GPT-5-Codex or Claude 4.5 Sonnet.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Code UI with Claude 4.5 Sonnet&lt;/strong&gt;. It&amp;rsquo;s better than most models.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;p&gt;Check out the &lt;a href=&#34;https://sanand0.github.io/datastories/browser-history/&#34;&gt;prompts &amp;amp; process&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Try this&lt;/strong&gt;: Pick one messy dataset you have. Ask an LLM for five ways to explore it. Build them all. One will surprise you.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.linkedin.com/posts/sanand0_last-weekend-i-fed-codex-my-browser-history-activity-7381531293542449152-uZXX&#34;&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title></title>
      <link>https://www.s-anand.net/blog/rip-data-scientists/</link>
      <pubDate>Fri, 12 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/rip-data-scientists/</guid>
      <description>&lt;p&gt;Slides for my DataHack Summit talk (controversially) titled &lt;strong&gt;RIP Data Scientists&lt;/strong&gt; are at &lt;a href=&#34;https://sanand0.github.io/talks/2025-08-21-rip-data-scientists/&#34;&gt;https://sanand0.github.io/talks/2025-08-21-rip-data-scientists/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Summary&lt;/strong&gt;: as data scientists we explore, clean, model, explain, deploy, and anonymize datasets. I live-vibe-coded &lt;em&gt;each&lt;/em&gt; step with DGCA data in 35 minutes using ChatGPT.&lt;/p&gt;
&lt;p&gt;Of course, it&amp;rsquo;s the &lt;em&gt;tasks&lt;/em&gt; that are dying, not the role. Data scientists will leverage AI, differentiate on other skills, and move on.&lt;/p&gt;
&lt;p&gt;But the highlight was an audience comment: &amp;ldquo;I&amp;rsquo;m no data scientist. I&amp;rsquo;m a domain person. I&amp;rsquo;ll tell you all this: If you don&amp;rsquo;t follow these practices, you won&amp;rsquo;t have a job with me!&amp;rdquo;&lt;/p&gt;
&lt;p&gt;They&amp;rsquo;re catching on! 🙂&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2025-09-12-rip-data-scientists-linkedin.jpg&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.linkedin.com/posts/sanand0_slides-for-my-datahack-summit-talk-controversially-activity-7364307983037579266-gnqf&#34;&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 17 Aug 2025</title>
      <link>https://www.s-anand.net/blog/things-i-learned-17-aug-2025/</link>
      <pubDate>Sun, 17 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-17-aug-2025/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Git &lt;a href=&#34;https://git-scm.com/docs/partial-clone&#34;&gt;partial clone&lt;/a&gt; lets you fetch files on-demand! E.g. &lt;code&gt;git clone --filter=&#39;blobs:size=100k&#39; &amp;lt;repo&amp;gt;&lt;/code&gt; will clone files under 100K and fetch the rest only on checkout. Over time, Git LFS capabilities will migrate into native Git. &lt;a href=&#34;https://tylercipriani.com/blog/2025/08/15/git-lfs/&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;⭐ From Daniel Kahneman, The Knowledge Project Podcast.
&lt;ul&gt;
&lt;li&gt;Key lesson. Have lower expectations. Behavior change is hard.&lt;/li&gt;
&lt;li&gt;Happiness is pleasure in the moment. Satisfaction is the meaningful story of our life. When reflecting, the thinking brain wants satisfaction. When feeling, the feeling brain feels happiness. The 2 brains optimize for different things. The thinking brain packs the calendar with satisfying tasks that the feeling brain hates doing.&lt;/li&gt;
&lt;li&gt;Happiness &amp;amp; pleasure are both are good for us. We don&amp;rsquo;t know which matters more.&lt;/li&gt;
&lt;li&gt;Behavior change is harder than most people think. Usually, it&amp;rsquo;s better not to expect success. Changing others, or ourselves.
&lt;ul&gt;
&lt;li&gt;Instead, &lt;em&gt;understand&lt;/em&gt; the cause of that behavior. Behaviour is an equilibrium of forces.&lt;/li&gt;
&lt;li&gt;Weakening forces preventing right behaviour is easier than strengthening forward forces. It lowers tension. That&amp;rsquo;s inversion!&lt;/li&gt;
&lt;li&gt;Behaviours are more about situations than personality. We assume otherwise - that&amp;rsquo;s an attribution error.&lt;/li&gt;
&lt;li&gt;Environment shapes thinking but it&amp;rsquo;s not obvious how, e.g. some people work better in noisy cafes. Some colors are more calming.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Leadership &amp;amp; delegation
&lt;ul&gt;
&lt;li&gt;Motivation is complex. People can do bad things for good reasons and vice versa.&lt;/li&gt;
&lt;li&gt;So, delegate decisions to unemotional agents. But agents misjudge perceived value of gain or loss!&lt;/li&gt;
&lt;li&gt;People prefer over-confident intuitive leaders over slow, deliberate leaders.&lt;/li&gt;
&lt;li&gt;Protect dissenters and dissent. It&amp;rsquo;s painful and costly, and needs nurturing.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Negotiation is about &lt;em&gt;understanding&lt;/em&gt;, not convincing.&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Feelings get in the way of clear thinking.&amp;rdquo;
&lt;ul&gt;
&lt;li&gt;Example: I vibe-coded the last 2 questions of &lt;a href=&#34;https://exam.sanand.workers.dev/tds-2025-05-ga7&#34;&gt;TDS GA7&lt;/a&gt; on Claude Code. It didn&amp;rsquo;t run. I delayed fixing it for 5 days, afraid it would a major effort. It ended up a 2 min fix. It &lt;em&gt;could&lt;/em&gt; have been major, but checking would have helped. Fear prevented that.&lt;/li&gt;
&lt;li&gt;Intuition, emotion, beliefs hamper clear thinking. Beliefs are often formed based on people we admire or identify, not reason.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;What enables clear thinking (all are hard):
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Pragmatism&lt;/strong&gt;. Don&amp;rsquo;t threaten your identity, the leader, etc. Else none of this works.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rules&lt;/strong&gt;, systems and processes. Willpower is illusion. Alignment is an illusion. &amp;ldquo;Whereever there is judgement, there is noise, and more than what people think.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Standards&lt;/strong&gt;. Shared, consistent scales of evaluation. Super-forecasters use probability scales.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deliberation&lt;/strong&gt;. Slow decision making.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Decomposition&lt;/strong&gt;. Break down the problem, analyze it, THEN form an intuition. Be disciplined in delaying intuition or forming an opinion.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pre-mortems&lt;/strong&gt;. &amp;ldquo;Write the history of the disaster this decision led to.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Decision journals&lt;/strong&gt; with post-mortems. Pros, cons and alternatives from failed decisions, e.g. Ray Dalio&amp;rsquo;s principles. Change of mind.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Independent data&lt;/strong&gt;. Use data. Keep evidence gatherers independent of decision makers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Preparation&lt;/strong&gt;. Have decision makers write down decisions &lt;em&gt;before&lt;/em&gt; discussing. Increases diversity.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;DuckDB&amp;rsquo;s feature engineering capabilites are faster than scikit-learn. &lt;a href=&#34;https://duckdb.org/2025/08/15/ml-data-preprocessing.html&#34;&gt;DuckDB&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Developers are encoding their &lt;em&gt;entire&lt;/em&gt; SDLC workflow into Claude commands &lt;a href=&#34;https://chatgpt.com/c/68a0139b-3044-8327-b2f0-51940f89b8ec&#34;&gt;ChatGPT&lt;/a&gt; #ai-coding
&lt;ul&gt;
&lt;li&gt;Commands are used for:
&lt;ul&gt;
&lt;li&gt;Requirements: Research sub-agent, task breakdown into todos.md, creating specs.md from todos.md&lt;/li&gt;
&lt;li&gt;Progress tracking: session logging, effort tracking, updating status, planning next steps&lt;/li&gt;
&lt;li&gt;Project setup: initializing, adding deps, scaffolding features&lt;/li&gt;
&lt;li&gt;Development: code review, debug error (five whys), explain code, refactor code&lt;/li&gt;
&lt;li&gt;Optimization: optimize build, DB, caching&lt;/li&gt;
&lt;li&gt;Testing: TDD, generate test cases, set up unit/integration/E2E testing, analyze coverage&lt;/li&gt;
&lt;li&gt;Security: security audits, dependency vulnerability scans&lt;/li&gt;
&lt;li&gt;Integration: sync tasks between GitHub and Linear (two-way issue synchronization, PR linking)&lt;/li&gt;
&lt;li&gt;Deployment: prepare releases, hotfix deploys, rollbacks, containerization, CI pipeline setup&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Patterns of usage
&lt;ul&gt;
&lt;li&gt;Sub-agents&lt;/li&gt;
&lt;li&gt;Command handoffs, i.e. one command invoking another&lt;/li&gt;
&lt;li&gt;Shared among a team in a repo, enforcing standards &amp;amp; sharing best practices&lt;/li&gt;
&lt;li&gt;Integration with specific tools / APIs (e.g. Linear)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;⭐ LLMs can hyper-personalize demos. E.g. an LLM document generator demo accepts a role, document type, and prompt. The demo-er says &amp;ldquo;Bank, LinkedIn marketing&amp;rdquo; and the LLM auto-populates the fields aptly, re-purposing the demo.&lt;/li&gt;
&lt;li&gt;From the &lt;a href=&#34;https://cdn.openai.com/API/docs/gpt-5-for-coding-cheatsheet.pdf&#34;&gt;GPT 5 coding cheatsheet&lt;/a&gt;:
&lt;ol&gt;
&lt;li&gt;Be precise and avoid conflicting information. Use a prompt optimizer to check for inconsistencies.&lt;/li&gt;
&lt;li&gt;Use the right reasoning effort. Prefer medium or low reasoning to avoid overthinking simple problems.&lt;/li&gt;
&lt;li&gt;Use XML-like syntax to help structure instructions&lt;/li&gt;
&lt;li&gt;Avoid overly firm language, e.g. &amp;ldquo;You MUST be THOROUGH&amp;rdquo; vs &amp;ldquo;Thoroughly&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;Give room for planning and self-reflection. Explain what to do in steps, asking it to think deeply&lt;/li&gt;
&lt;li&gt;Control the eagerness of your coding agent, e.g. do not ask for confirmation, parallelize tool calls, use more tools, etc.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;⭐ Assets are any leveragable stored capability. Money is one, but there are several one can &amp;ldquo;invest&amp;rdquo; in, be an agent of, or perhaps steal.
&lt;ol&gt;
&lt;li&gt;Wealth (investments, income)&lt;/li&gt;
&lt;li&gt;Regenerative assets (land, carbon credits, renewables)&lt;/li&gt;
&lt;li&gt;Contacts (reference customers, hiring pipeline, talent bench, weak-ties)&lt;/li&gt;
&lt;li&gt;Distribution channels (repeatable routes to users: partnerships, marketplaces, APIs, SEO)&lt;/li&gt;
&lt;li&gt;Attention (your audience, whom you can reach directly)&lt;/li&gt;
&lt;li&gt;Trust/reputation in communities (community capital in employers, clients, forums, society, search keywords)&lt;/li&gt;
&lt;li&gt;Personal brand “edges” (moral authority, values lived aloud, distinctive taste or stance)&lt;/li&gt;
&lt;li&gt;Data (your clean, labeled, joined data corpus)&lt;/li&gt;
&lt;li&gt;Code (models, algorithms, components, templates, libraries, tools, evals; versioned)&lt;/li&gt;
&lt;li&gt;Content (blog posts, video tutorials, case studies, demos, stories, slides, docs)&lt;/li&gt;
&lt;li&gt;Knowledge (notes, decision logs, knowledge graph, institutional memory)&lt;/li&gt;
&lt;li&gt;Playbooks &amp;amp; runbooks (process checklists that survived fire, SOPs, scenario plans)&lt;/li&gt;
&lt;li&gt;Habits &amp;amp; policies (operating cadence, rituals, governance &amp;amp; compliance muscle)&lt;/li&gt;
&lt;li&gt;Optionality (cash buffer, credit lines, slack time, real options, small bets)&lt;/li&gt;
&lt;li&gt;Agreements (MSAs/SLAs, pre-negotiated contracts)&lt;/li&gt;
&lt;li&gt;IP (copyrights, trade secrets, trademarks)&lt;/li&gt;
&lt;li&gt;Health &amp;amp; energy reserves&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;⭐ Intense negative emotions get in the way of clear thinking. Curiosity, humor, kindness, and gratitude help. (Intense positive emotions like awe, passion, etc. help creativity and are not so bad.) #beliefs&lt;/li&gt;
&lt;li&gt;I like to think I&amp;rsquo;m a Python expert. When I saw a client use this code, I told her the indentation is wrong. It ran just fine. And people think only LLMs hallucinate.&lt;/li&gt;
&lt;li&gt;This is undocumented, but the way to get an &lt;a href=&#34;https://ai.google.dev/api/live#ephemeral-auth-tokens&#34;&gt;Gemini ephemeral auth token&lt;/a&gt; for the live API is below. (Update time as required.) &lt;a href=&#34;https://chatgpt.com/share/689f591e-aa08-800c-b272-dba3abe1ee37&#34;&gt;ChatGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Learnings from a discussion on vibe-coding between &lt;a href=&#34;https://www.linkedin.com/in/jaink/&#34;&gt;Kunal Jain&lt;/a&gt;, &lt;a href=&#34;https://www.linkedin.com/in/ever-loyal/&#34;&gt;Ravi Nadimpalli&lt;/a&gt; and me. #ai-coding
&lt;ul&gt;
&lt;li&gt;On the Vibe Coding Process &amp;amp; Strategy
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The 80/20 Rule is Real:&lt;/strong&gt; The first 80% of a project is incredibly fast, but the final 20% (debugging, custom features, production-readiness) is extremely difficult and time-consuming.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Validation is the New Bottleneck:&lt;/strong&gt; Since coding is now much faster, the critical, time-consuming task has shifted to reviewing, testing, and validating the LLM&amp;rsquo;s output.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&amp;ldquo;Spec-Locking&amp;rdquo; is Crucial:&lt;/strong&gt; Providing the LLM with detailed, well-defined, and &amp;ldquo;thinly sliced&amp;rdquo; specifications is essential for getting good results. Vague requests lead to poor outcomes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It&amp;rsquo;s Not Production-Ready (Yet):&lt;/strong&gt; The consensus is that vibe coding is excellent for prototypes, demos, and go-to-market (GTM) activities but is not yet reliable for building production-grade applications from scratch.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Code is Brittle &amp;amp; Unstable:&lt;/strong&gt; An application that works perfectly one day can inexplicably break the next, as the underlying agent might make undocumented changes.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Impact on Roles &amp;amp; The Future of Work
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;The Rise of QC/Validation:&lt;/strong&gt; The Quality Control (QC) function will become larger and more critical to manage the new challenge of validating AI-generated work.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Product Managers Shift Focus:&lt;/strong&gt; PMs can move away from tedious documentation (like flowcharts) and focus more on high-level business strategy, using vibe coding to create quick prototypes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Democratization of Building:&lt;/strong&gt; It empowers non-coders to build functional apps and helps professionals upskill faster by &amp;ldquo;conversing&amp;rdquo; with an LLM on complex topics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;New Forms of Cheating:&lt;/strong&gt; The technology is creating novel ways for people to cheat in interviews, such as using tools that provide real-time subtitles of answers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The &amp;ldquo;Jagged Edge&amp;rdquo; of AI:&lt;/strong&gt; The technology excels at certain tasks (like GTM content) but fails at others, creating new upstream bottlenecks where teams must rapidly generate more of the &amp;ldquo;AI-friendly&amp;rdquo; work.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Practical Hacks &amp;amp; Takeaways
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Meta-Prompting:&lt;/strong&gt; Use an LLM to refine and improve your prompt &lt;em&gt;before&lt;/em&gt; giving it to the final tool. This helps fill in gaps and add necessary detail.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Human-First Drafting:&lt;/strong&gt; For creative or nuanced work (like writing), it&amp;rsquo;s often better to write the first draft yourself and use the LLM to polish it, rather than starting with a generic AI draft.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Use Structured Prompts:&lt;/strong&gt; For predictable and clean output, providing instructions in a structured format (JSON is OK but not needed) is highly effective.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM as a Judge:&lt;/strong&gt; Use LLMs to evaluate and grade content, code, and other outputs, dramatically speeding up the review process.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Automate Learning &amp;amp; Documentation:&lt;/strong&gt; Use tools to transcribe conversations automatically and create personalized revision quizzes from notes and documents.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Voice is a Powerful Modality:&lt;/strong&gt; Using voice-to-code allows for capturing more complex ideas faster and can be done while multitasking (e.g., walking), capitalizing on &amp;ldquo;dead time.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;For live transcription, Gemini 2.5 Flash Live costs 0.6c/min of audio ($3/MTok x 32 tokens/second) while GPT 4o Mini Realtime costs ~2c/min and GPT 4o Realtime costs ~8c/min. &lt;a href=&#34;https://chatgpt.com/share/689ef64f-2510-800c-83c0-052bbbf28acf&#34;&gt;ChatGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;I set up MCPs &lt;a href=&#34;https://github.com/openai/codex&#34;&gt;Codex CLI&lt;/a&gt; by adding this to &lt;code&gt;~/.codex/config.toml&lt;/code&gt;. I&amp;rsquo;ve disabled it for faster startup (this takes ~2 seconds) and raised an enhancement &lt;a href=&#34;https://github.com/openai/codex/issues/2335&#34;&gt;issue for MCP lazy loading&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.anthropic.com/news/agent-capabilities-api&#34;&gt;Anthropic&lt;/a&gt; launched a remote MCP connector in their API. OpenAI Responses API &lt;a href=&#34;https://platform.openai.com/docs/guides/tools-remote-mcp&#34;&gt;already had remote MCP support&lt;/a&gt;. Gemini will likely follow, opening up new tool capabilities. The APIs can &lt;em&gt;directly&lt;/em&gt; call the MCPs as part of their thinking.&lt;/li&gt;
&lt;li&gt;Turns out Indian English is a well studied topic. Indianisms like &amp;ldquo;can able to&amp;rdquo;, &amp;ldquo;need not to&amp;rdquo;, &amp;ldquo;why because…&amp;rdquo;, &amp;ldquo;if suppose…&amp;rdquo;, &amp;ldquo;return back&amp;rdquo;, &amp;ldquo;revert back&amp;rdquo;, &amp;ldquo;angry on&amp;rdquo;, &amp;ldquo;discuss about&amp;rdquo;, &amp;ldquo;order for&amp;rdquo;, &amp;ldquo;do one thing…&amp;rdquo;, &amp;ldquo;give me a missed call&amp;rdquo;, &amp;ldquo;what is your good name&amp;rdquo;, &amp;ldquo;kindly adjust&amp;rdquo;, &amp;ldquo;we are like that only&amp;rdquo;, &amp;ldquo;he is coming only&amp;rdquo;, &amp;ldquo;today itself&amp;rdquo;, &amp;ldquo;now only&amp;rdquo;, &amp;ldquo;prepone&amp;rdquo;, &amp;ldquo;pass out (of college)&amp;rdquo;, &amp;ldquo;out of station&amp;rdquo;, &amp;ldquo;do the needful&amp;rdquo;, &amp;ldquo;hotel&amp;rdquo;, &amp;ldquo;batchmate&amp;rdquo;, &amp;ldquo;cousin-brother / cousin-sister&amp;rdquo;, &amp;ldquo;I have a doubt&amp;rdquo;, &amp;ldquo;I am understanding&amp;rdquo;, &amp;ldquo;she is knowing&amp;rdquo;, &amp;ldquo;you’re coming, no?&amp;rdquo; etc. are discussed in &lt;a href=&#34;https://theswissbay.ch/pdf/Books/Linguistics/Mega%20linguistics%20pack/Indo-European/Germanic/English%2C%20Indian%20%28Sailaja%29.pdf&#34;&gt;Pingali Sailaja&amp;rsquo;s Indian English&lt;/a&gt;. &lt;a href=&#34;https://chatgpt.com/share/689dcf8d-2ce4-800c-8553-e419eafd4891&#34;&gt;ChatGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Astral is &lt;a href=&#34;https://astral.sh/blog/introducing-pyx&#34;&gt;building pyx&lt;/a&gt; - a paid PyPi alternative. It aims to solve problems like PyTorch CUDA builds. Knowing them, it&amp;rsquo;ll be fabulous. I look forward to when they build a Python hosting service.&lt;/li&gt;
&lt;li&gt;⭐ Here&amp;rsquo;s one way to improve LLMs apps in real-time.
&lt;ul&gt;
&lt;li&gt;After sending a response, send the prompt + input + output + optional user feedback to an LLM-as-a-judge asking for feedback to improve the prompt.&lt;/li&gt;
&lt;li&gt;Revise the prompt based on the improvement. Now the app has improved, real-time, based on human/LLM feedback.&lt;/li&gt;
&lt;li&gt;Refine this process to ensure that the revisions are smooth and positive.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;GPT 4.1 (and presumably GPT 5) models have been trained on a &lt;a href=&#34;https://cookbook.openai.com/examples/gpt4-1_prompting_guide#appendix-generating-and-applying-file-diffs&#34;&gt;specific diff format&lt;/a&gt; useful for code diff-patching. &lt;a href=&#34;https://github.com/12458/PseudoPatch&#34;&gt;PseudoPatch&lt;/a&gt; is a Python package that implements their &lt;code&gt;apply_patch()&lt;/code&gt; function. Aider supports multiple &lt;a href=&#34;https://aider.chat/docs/more/edit-formats.html&#34;&gt;edit formats&lt;/a&gt; that are commonly referenced as a standard. &lt;a href=&#34;https://fabianhertwig.com/blog/coding-assistants-file-edits/&#34;&gt;Code Surgery&lt;/a&gt; has a good walkthrough of various strategies. These are similar to Google&amp;rsquo;s &lt;a href=&#34;https://github.com/google/diff-match-patch&#34;&gt;diff-match-patch&lt;/a&gt; approach (which fuzzy matches and &lt;em&gt;then&lt;/em&gt; patches) but does not require line numbers. &lt;a href=&#34;https://chatgpt.com/share/689753f9-7b24-800c-b568-4ff8c7978486&#34;&gt;ChatGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Here are some query parameters &lt;a href=&#34;https://chatgpt.com/&#34;&gt;ChatGPT.com&lt;/a&gt; unofficially supports:
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;?q=...&lt;/code&gt; prefills in a new chat &lt;strong&gt;and often auto-submits&lt;/strong&gt;, especially small text &lt;a href=&#34;https://treyhunner.com/2024/07/chatgpt-and-claude-from-your-browser-url-bar/&#34;&gt;#&lt;/a&gt;. Useful for:
&lt;ul&gt;
&lt;li&gt;A custom search engine in your browser&lt;/li&gt;
&lt;li&gt;An &amp;ldquo;Ask ChatGPT about selection&amp;rdquo; bookmarklet, etc.&lt;/li&gt;
&lt;li&gt;Links (e.g. from courses, FAQs, etc.) for tasks or learning&lt;/li&gt;
&lt;li&gt;&amp;hellip; but not for custom GPTs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;?model=...&lt;/code&gt; selects a model (e.g., &lt;code&gt;gpt-5-thinking&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;&lt;code&gt;?hints=search&lt;/code&gt; enables &lt;strong&gt;Search&lt;/strong&gt; mode&lt;/li&gt;
&lt;li&gt;&lt;code&gt;?temporary-chat=true&lt;/code&gt; opens a new temporary chat&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.tavus.io/&#34;&gt;Tavus&lt;/a&gt; is another AI avatar platform.
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;&#34;&gt;Synthesia&lt;/a&gt;. Market leader; $2.1B valuation; enterprise trusted. Good: Realism, enterprise features, templating. But: Price, usage caps, slower avatar setup&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;&#34;&gt;HeyGen&lt;/a&gt;. Rapidly growing; $500M valuation. Good: Avatar realism, speed, affordability. But: Basic collaboration, support, scene complexity&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;&#34;&gt;Colossyan&lt;/a&gt;. Favored L&amp;amp;D focus. Good: Interactive &amp;amp; educational tools, good value. But: Less polished avatars, slower renders&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;&#34;&gt;D-ID&lt;/a&gt;. Frequently cited alternative. Good: Speed, flexibility, custom avatars. But: Watermarks, fewer templates&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;&#34;&gt;Elai.io&lt;/a&gt;. Repeats in alternatives lists. Good: Storyboarding, educational formats. But: Limited templates, render time&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;&#34;&gt;Hour One&lt;/a&gt;. Also common in alternative lists. Good: Photoreal avatars, expression control. But: Missing advanced features like screen capture&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;&#34;&gt;Others&lt;/a&gt;. Niche or emerging tools. Good: Varies by platform. But: Less adoption, fewer reviews&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Training companies are offering &amp;ldquo;Labs-as-a-service&amp;rdquo; as part of their AI training. Corporates ban LLMs, but need employees trained. Trainers offer a bundled package where they also offer access to LLMs are part of their course. Interesting business-model value-add.&lt;/li&gt;
&lt;li&gt;⭐ I&amp;rsquo;m meta-AI-coding. I wrote a crude prompt in &lt;code&gt;prompts.md&lt;/code&gt;, told Codex &amp;ldquo;prompts.md has a prompt under the &amp;ldquo;# Improve schema&amp;rdquo; section starting line 294. This is a prompt that will be passed to Claude Code to implement. Ask me questions as required and improve the prompt so that the results will be in line with my expectations, one-shot.&amp;rdquo; After a few discussions, it generated &lt;a href=&#34;https://github.com/sanand0/slidegen/blame/de953817266357b00d80d4fa3e17def02e0de292/prompts.md#L296-L502&#34;&gt;this remarkable prompt&lt;/a&gt;. This prompt was easy for me to review AND easy for Claude Code to understand because of the lack of inconsistencies.
&lt;ul&gt;
&lt;li&gt;Use the &lt;strong&gt;Ask-Code pattern&lt;/strong&gt;. In Codex, speak the requirement and have it rewrite the prompt asking clarifying questions &lt;em&gt;pressing the &lt;strong&gt;Ask&lt;/strong&gt; button&lt;/em&gt; instead of Code. Then, answer its questions. &lt;em&gt;Then&lt;/em&gt; press &lt;strong&gt;Code&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A Forward Deployed Engineer (FDE) is a hybrid role, part software engineer, part product manager, and part consultant, focused on deeply integrating a company&amp;rsquo;s technology with a specific client&amp;rsquo;s needs.&lt;/li&gt;
&lt;li&gt;Based on what I&amp;rsquo;ve seen of AI coding, new developers need to learn these skills. #ai-coding
&lt;ul&gt;
&lt;li&gt;Context engineering&lt;/li&gt;
&lt;li&gt;Documentation&lt;/li&gt;
&lt;li&gt;Automated testing&lt;/li&gt;
&lt;li&gt;Standards&lt;/li&gt;
&lt;li&gt;Capabilities of platforms&lt;/li&gt;
&lt;li&gt;Modularity (and DRY vs WET)&lt;/li&gt;
&lt;li&gt;Code composition&lt;/li&gt;
&lt;li&gt;Code reviews&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Blindspots continue to be the insight with maximum RoI. Discovering something we&amp;rsquo;re not even aware we&amp;rsquo;re unaware of opens up the largest possibilities. #beliefs My top sources to discover blindspots are:
&lt;ul&gt;
&lt;li&gt;Feedback. Especially feedback we reject, ignore, or miss.&lt;/li&gt;
&lt;li&gt;Things we run/shy away from.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Across clients, providers (e.g. Bedrock) and products (e.g. Cursor) I have observed capacity bottlenecks for Claude models which don&amp;rsquo;t seem to affect OpenAI models as much.&lt;/li&gt;
&lt;li&gt;Increasing the size of an image improves OCR accuracy for LLM models (or at least Claude 4 Sonnet). Anecdotally, resizing 2x did not work on a number of examples but 2.5x - 3x did. This increases the cost to 6.25x or 9x, however.&lt;/li&gt;
&lt;li&gt;Discussion at &lt;a href=&#34;https://pycon.sg/edusummit.html&#34;&gt;PyConSG Edu Summit 2025&lt;/a&gt;. &lt;a href=&#34;https://padlet.com/pyconsg/pyconsg-education-summit-2025-topic-how-to-prevent-students--57puwelj2o7rgadd&#34;&gt;Padlet&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://chatgpt.com/share/6899bd13-03e0-800c-8618-971ed7050a1a&#34;&gt;Discussion validation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Interesting ways students use AI
&lt;ul&gt;
&lt;li&gt;Use AI to refactor/debug whole codebases&lt;/li&gt;
&lt;li&gt;Get AI to create questions for practice&lt;/li&gt;
&lt;li&gt;ChatGPT Study mode&lt;/li&gt;
&lt;li&gt;Students like to upload photos. We can teach them to upload these to ChatGPT and ask questions.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;What teaching practices / assessment design can help students think for themselves before turning to AI? &lt;a href=&#34;https://chatgpt.com/share/6899bc2c-4678-800c-b133-3653c378e978&#34;&gt;ChatGPT&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Interactive orals / micro-vivas (short, process-focused).&lt;/strong&gt; Strong alignment with “interactive oral assessment” research and guidance in the AI era: improves authenticity, reduces outsourcing/contract cheating, and checks understanding. Make them low-stakes but frequent.
&lt;em&gt;How&lt;/em&gt;: 5–8 min viva tied to a task; students must explain choices, failures, and next steps.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Authentic / project-based assessments students can self-validate (observable outputs).&lt;/strong&gt; Project-based and “authentic” assessment meta-reviews show consistent positive effects (achievement, thinking skills, motivation), especially in STEM and small teams. Design tasks with &lt;em&gt;local data/constraints&lt;/em&gt; so generic LLM answers are only a baseline.
&lt;em&gt;How&lt;/em&gt;: “Default AI answer” gets a pass; “A-grade” requires empirical validation, custom data, or optimisation trade-offs with metrics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pair programming + peer critique on whiteboards/pseudocode.&lt;/strong&gt; Evidence (meta-analyses &amp;amp; CS-ed studies) supports pair programming for learning and retention; code tracing/peer instruction deepen understanding before coding.
&lt;em&gt;How&lt;/em&gt;: Rotate driver/navigator; force commit-message style rationales; 10-minute “whiteboard dry-run” before touching IDE.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Process-over-product with structured reflection.&lt;/strong&gt; Metacognitive/reflective interventions show medium-to-large effects on achievement; they also build habits that resist blind acceptance of AI outputs. Keep reflections short but structured.
&lt;em&gt;How&lt;/em&gt;: “What I asked AI; what it missed; how I verified; what I’d change next time.”&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;“No-AI under secure conditions” mixed with AI-permitted coursework.&lt;/strong&gt; Matches national/institutional guidance for GenAI-aware assessment design. Use secure, time-boxed checks for fundamentals; allow AI elsewhere with audit trails.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Primary research (interviews/user studies) before design/coding.&lt;/strong&gt; Fits the “authentic assessment” literature and reduces LLM substitution. Grade on research protocol + synthesis rigor, not word count.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Explicit problem-solving frames (initial/current/goal state).&lt;/strong&gt; Classic problem-solving scaffolds; improves formulation before querying AI. Pair with short “assumption logs.” (General pedagogy supported; CT depends on domain knowledge &amp;ndash; see caveat below.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Caveat (important):&lt;/strong&gt; &lt;em&gt;Critical thinking depends on domain knowledge.&lt;/em&gt; Don’t expect generic CT drills to transfer without content mastery. Plan tasks so students must recall/apply &lt;em&gt;specific&lt;/em&gt; knowledge before or alongside AI.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;How can we train students to use AI critically instead of accepting the output blindly? &lt;a href=&#34;https://chatgpt.com/share/6899bc5e-1800-800c-bfce-25d261c63a09&#34;&gt;ChatGPT&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Teach “lateral reading” and SIFT for source checking.&lt;/strong&gt; Stanford’s Civic Online Reasoning work and Caulfield’s SIFT method offer actionable heuristics for verifying claims, URLs, and citations that LLMs surface. Build these into rubrics.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Run “AI auditing” labs (hallucination hunts).&lt;/strong&gt; Students collect/label model mistakes, missing assumptions, and fabricated citations &amp;ndash; an approach aligned with UNESCO’s call for AI literacy and validation.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Use online judges with &lt;em&gt;hidden&lt;/em&gt; tests + adversarial cases.&lt;/strong&gt; Autograding literature supports hidden tests for robust generalization; it trains students to verify and not overfit to visible specs &amp;ndash; or to AI’s surface patterns.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;“Sandwich” workflow: spec → implement 1–2 reps → let AI complete → &lt;em&gt;verify&lt;/em&gt; rigorously.&lt;/strong&gt; Mirrors human-in-the-loop patterns in industry; use checklists for unit/property tests and invariants before accepting AI output.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Live-coding with an AI assistant &lt;em&gt;on display&lt;/em&gt; (to show failure modes).&lt;/strong&gt; Demonstrates nondeterminism/limitations in real time; supports critical habits. Pair with a post-mortem template.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prompt red-teaming/jailbreak exercises (safe scope).&lt;/strong&gt; Students learn that guardrails can be bypassed and why verification matters. Keep it ethical and bounded.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Build a knowledge base first.&lt;/strong&gt; Reinforce that CT sits on content knowledge; teach students to &lt;em&gt;explain&lt;/em&gt; why an AI answer is plausible or not, citing domain facts.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes from &amp;ldquo;My Thoughts on Computational Thinking in the Generative AI Era&amp;rdquo; by &lt;a href=&#34;https://www.comp.nus.edu.sg/cs/people/leonghw/&#34;&gt;LEONG Hon Wai&lt;/a&gt;, ex-NUS, at &lt;a href=&#34;https://pycon.sg/edusummit.html&#34;&gt;PyConSG Edu Summit 2025&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Students from China don&amp;rsquo;t like to write, express their ideas, and share. That&amp;rsquo;s changing now.&lt;/li&gt;
&lt;li&gt;Computational thinking is pretty new (Jeannette Wing, 2006), actually, based on Papert (1980). It&amp;rsquo;s too early to abandon it.&lt;/li&gt;
&lt;li&gt;It enables effective learning attitudes:
&lt;ul&gt;
&lt;li&gt;Tinker (experiment &amp;amp; play): helps finding diverse problems to generalize into&lt;/li&gt;
&lt;li&gt;Debug (find &amp;amp; fix bugs)&lt;/li&gt;
&lt;li&gt;Create (design &amp;amp; make)&lt;/li&gt;
&lt;li&gt;Persevere (keep going): but only if it&amp;rsquo;s &lt;em&gt;productive&lt;/em&gt;, i.e failing in &lt;em&gt;new&lt;/em&gt; ways&lt;/li&gt;
&lt;li&gt;Collaborate &amp;amp; communicate&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Teaching this is hard. Get students to &lt;em&gt;WANT&lt;/em&gt; to do computational thinking.&lt;/li&gt;
&lt;li&gt;Problem formulation (among the computational thinking blocks) is more important than before.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://cacm.acm.org/blogcacm/leveraging-computational-thinking-in-the-era-of-generative-ai/&#34;&gt;Leveraging Computational Thinking in the Era of Generative AI&lt;/a&gt; argues that computational thinking manifests in prompt/context engineering.&lt;/li&gt;
&lt;li&gt;We&amp;rsquo;re moving from &amp;ldquo;Computational Thinking&amp;rdquo; to &amp;ldquo;Computational Action&amp;rdquo; &amp;ndash; where we&amp;rsquo;re talking to AI coders that actually deploy apps that &lt;em&gt;DO&lt;/em&gt; stuff.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes from &amp;ldquo;Make Learning Easy and Fun @ NLB LearnX&amp;rdquo; by Goh Soon Seng, NLB, at &lt;a href=&#34;https://pycon.sg/edusummit.html&#34;&gt;PyConSG Edu Summit 2025&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Libraries have a Pi Python Makers Club, open for all. Bi-monthly meetings. Quarterly Pi Python workshop.&lt;/li&gt;
&lt;li&gt;Space provides 3D printers, Raspberry Pi, sensors, etc.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes from &amp;ldquo;Teaching Goals and Plans - How we might help students improve problem-solving&amp;rdquo; by Dr Norman Lee, SUTD, at &lt;a href=&#34;https://pycon.sg/edusummit.html&#34;&gt;PyConSG Edu Summit 2025&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Programming is &lt;em&gt;hard&lt;/em&gt;. E.g. Solving the &lt;a href=&#34;https://scholar.google.com/scholar?q=Soloway+rainfall+problem&#34;&gt;Rainfall problem&lt;/a&gt; &amp;ldquo;Sum numbers until 99999&amp;rdquo; needs &lt;em&gt;several&lt;/em&gt; building blocks:
&lt;ul&gt;
&lt;li&gt;Python syntax&lt;/li&gt;
&lt;li&gt;Getting user input&lt;/li&gt;
&lt;li&gt;While loop&lt;/li&gt;
&lt;li&gt;Controlling while loop with counter&lt;/li&gt;
&lt;li&gt;Accumulation&lt;/li&gt;
&lt;li&gt;If-else&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Merging (or composing) such blocks is the hard part. In &lt;a href=&#34;https://scholar.google.com/scholar?cluster=16826723591053220162&#34;&gt;Learning to program = learning to construct mechanisms and explanations&lt;/a&gt;, Soloway, shares 4 compositions.
&lt;ul&gt;
&lt;li&gt;Abutment: Put one block &lt;em&gt;after&lt;/em&gt; another&lt;/li&gt;
&lt;li&gt;Nesting: Put one block &lt;em&gt;inside&lt;/em&gt; another&lt;/li&gt;
&lt;li&gt;Merging: Interleave the code in the blocks&lt;/li&gt;
&lt;li&gt;Tailoring: Modify the code in the blocks&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;But you need to already have those primitives (patterns) to put together. The &amp;ldquo;expert blind spot&amp;rdquo; blinds experts to this.&lt;/li&gt;
&lt;li&gt;Actionable ideas:
&lt;ol&gt;
&lt;li&gt;Teach &lt;em&gt;patterns&lt;/em&gt; explicitly&lt;/li&gt;
&lt;li&gt;Create exercises on &lt;em&gt;applying&lt;/em&gt; them&lt;/li&gt;
&lt;li&gt;Use &lt;a href=&#34;https://en.wikipedia.org/wiki/Parsons_problem&#34;&gt;Parsons problem&lt;/a&gt;s: Fill in the blanks. Re-order lines of code. &lt;strong&gt;But&lt;/strong&gt; design problem carefully&lt;/li&gt;
&lt;li&gt;Step through a debugger. &lt;strong&gt;BUT&lt;/strong&gt; students must predict next line, not passive watching&lt;/li&gt;
&lt;li&gt;Teach to from one format (psuedocode, flowchart, another language like Excel) to Python. Helps multiple modes of learning&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes from &amp;ldquo;AISG programmes&amp;rdquo; by Chen Qeiquang, AI Singapore, &lt;a href=&#34;https://aiap.sg/apprenticeship/&#34;&gt;AI Apprentice Programme (AIAP)&lt;/a&gt; Assistant Head
&lt;ul&gt;
&lt;li&gt;Full-time. For SG citizens. $4,000/month. Build 3-6 month MVPs for startups, SMEs, or corporates. 300/1000 delivered so far.&lt;/li&gt;
&lt;li&gt;No lectures/tutorials. Focus is: topic assignments, discussion with mentors, apprentice sharing sessions.&lt;/li&gt;
&lt;li&gt;Includes an &lt;a href=&#34;https://aiap.sg/ladp/&#34;&gt;LLM Application Developer Program&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes from &amp;ldquo;Scaffolding the Problem-Solving Process for Introductory Computing Students&amp;rdquo; by Ashish Dandekar, NUS, at &lt;a href=&#34;https://pycon.sg/edusummit.html&#34;&gt;PyConSG Edu Summit 2025&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://scholar.google.com/scholar?cluster=5380873998289933948&#34;&gt;Built an intelligent tutoring system&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Encourage students to create their own pattern banks / cheat sheets. &amp;ldquo;Find 2 more problems that can be solved in the same way.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Focusing on the problem-solving process &lt;strong&gt;shrinks&lt;/strong&gt; the gap. Students &lt;em&gt;above&lt;/em&gt; the 50th percentile of pre-assessment did not improve much. The lowest percentile improved the most.&lt;/li&gt;
&lt;li&gt;&amp;ldquo;At NUS, I know that even if I give 0.5% weightage for students attending tutorials, &lt;em&gt;everyone&lt;/em&gt; will attend it for those &amp;lsquo;free marks&amp;rsquo;.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes from &amp;ldquo;Exploring Multi-Agent Generative AI in Education and Career Advisory&amp;rdquo; by Dr Yeo Wee Kiang, NUS, at &lt;a href=&#34;https://pycon.sg/edusummit.html&#34;&gt;PyConSG Edu Summit 2025&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;⭐ &amp;ldquo;When you have a high fever, do you speak more sense or nonsense? Nonsense. LLM temperature is like that. But it can also sound creative!&amp;rdquo;&lt;/li&gt;
&lt;li&gt;The router pattern is a powerful query rewriter. Redirects the query to specialized prompts/agents.&lt;/li&gt;
&lt;li&gt;Useful tools you can build for students: Course Mentor, Interview Coach, Job planner/matcher.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes from &amp;ldquo;Do we need to teach coding given vibe-coding tools?&amp;rdquo; by Dr. Oka Kurniawan, SUTD, at &lt;a href=&#34;https://pycon.sg/edusummit.html&#34;&gt;PyConSG Edu Summit 2025&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Paper: &lt;a href=&#34;https://journals.sagepub.com/doi/pdf/10.1177/15291006241287726&#34;&gt;What the Science of Learning Teaches Us About Arithmetic Fluency&lt;/a&gt; says mental math helps mathematicians. Fluency bootstraps higher-level thinking.&lt;/li&gt;
&lt;li&gt;MIT Media Lab&amp;rsquo;s Project: &lt;a href=&#34;https://www.media.mit.edu/projects/your-brain-on-chatgpt/overview/&#34;&gt;Your Brain on ChatGPT&lt;/a&gt;. Explores impact on brain. Bran-only group had the widest ranging brain networks. AI accumulates &lt;strong&gt;cognitive debt&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Paper: &amp;ldquo;A Study of the Difficulties of Novice Programmers&amp;rdquo; struggle with:
&lt;ol&gt;
&lt;li&gt;Syntax&lt;/li&gt;
&lt;li&gt;Problem solving&lt;/li&gt;
&lt;li&gt;Tools&lt;/li&gt;
&lt;li&gt;Computing concepts&lt;/li&gt;
&lt;li&gt;Analytical thinking / debugging&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Polya&amp;rsquo;s &lt;a href=&#34;https://en.wikipedia.org/wiki/How_to_Solve_It&#34;&gt;How to Solve It&lt;/a&gt; is the base problem solving framework for maths and can be adapted to computing&lt;/li&gt;
&lt;li&gt;Expert programmers have enough patterns to match against. Novices don&amp;rsquo;t. We need a &lt;strong&gt;bottoms-up framework&lt;/strong&gt; instead
&lt;ul&gt;
&lt;li&gt;Give them a concrete case.&lt;/li&gt;
&lt;li&gt;Have them generalize (loops, functional, vectors)&lt;/li&gt;
&lt;li&gt;Have them implement (debugging)&lt;/li&gt;
&lt;li&gt;Have them break it (test)&lt;/li&gt;
&lt;li&gt;All via &lt;strong&gt;vibe-coding&lt;/strong&gt;!&lt;/li&gt;
&lt;li&gt;The chats are tracked!!&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Paper: &lt;a href=&#34;https://researchrepository.ucd.ie/rest/bitstreams/41008/retrieve&#34;&gt;First Things First: Providing Metacognitive Scaffolding for Interpreting Problem Prompts&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Students often get the problem wrong&lt;/li&gt;
&lt;li&gt;Reading student conversations helps figure it out&lt;/li&gt;
&lt;li&gt;LLMs can figure it out too!&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Paper: &lt;a href=&#34;https://dl.acm.org/doi/pdf/10.1145/3632620.3671116&#34;&gt;The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Good coders got better with AI. Were able to ignore unhelpful advice.&lt;/li&gt;
&lt;li&gt;Poor coders got &lt;strong&gt;worse&lt;/strong&gt;! Thought they performed better than they did. &lt;em&gt;Increased&lt;/em&gt; illusion of competence.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The &lt;a href=&#34;https://challenge.bebraschallenge.org/&#34;&gt;Bebras Challenge&lt;/a&gt; is a global non-programming computational thinking (CT) challenge. &lt;a href=&#34;https://www.bebras.org/task-examples&#34;&gt;Examples&lt;/a&gt;. Singapore runs a &lt;a href=&#34;https://simcc.org/njio/&#34;&gt;National Junior Informatics Olympiad&lt;/a&gt; that learns from Bebras. It tests the &lt;em&gt;mindset&lt;/em&gt; behind coding, specifically &amp;ldquo;computational thinking&amp;rdquo;:
&lt;ul&gt;
&lt;li&gt;Problem formulation (added recently, and is increasingly important)&lt;/li&gt;
&lt;li&gt;Decomposition (and composition): break the problem down&lt;/li&gt;
&lt;li&gt;Pattern recognition: find the building blocks&lt;/li&gt;
&lt;li&gt;Abstraction: generalize useful blocks, drop irrelevant ones&lt;/li&gt;
&lt;li&gt;Algorithmic thinking: write the steps to solve&lt;/li&gt;
&lt;li&gt;Validation (not part of original list, but critical): how to efficiently check if this works&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://apple.github.io/embedding-atlas/&#34;&gt;Apple&amp;rsquo;s Embedding Atlas&lt;/a&gt; (&lt;a href=&#34;https://apple.github.io/embedding-atlas/demo/index.html&#34;&gt;Demo&lt;/a&gt; - slow, needs WebGPU) is an embeddings visualizer, like
&lt;a href=&#34;https://projector.tensorflow.org/&#34;&gt;Tensorflow Projector&lt;/a&gt; or &lt;a href=&#34;https://home.withmantis.com/&#34;&gt;Mantis&lt;/a&gt; (&lt;a href=&#34;https://mantisdev.csail.mit.edu/home/&#34;&gt;Demo&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;John Kotter&amp;rsquo;s organizational change model is the accepted practice for top-down change, while ADKAR is for bottom up. It&amp;rsquo;s surprising how obviously effective both are to someone who has effected both kinds of changes, but there is NO WAY I would have appreciated either during my MBA. &lt;a href=&#34;https://en.wikipedia.org/wiki/Change_management&#34;&gt;Wikipedia: Change management&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The OpenAI Chat Completions API has a few interesting and (relatively) new options:
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.openai.com/docs/api-reference/chat/create#chat_create-verbosity&#34;&gt;&lt;code&gt;verbosity&lt;/code&gt;&lt;/a&gt;. &lt;code&gt;low&lt;/code&gt;: concise response, &lt;code&gt;medium&lt;/code&gt;: default, &lt;code&gt;high&lt;/code&gt;: verbose&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.openai.com/docs/api-reference/chat/create#chat_create-reasoning_effort&#34;&gt;&lt;code&gt;reasoning_effort&lt;/code&gt;&lt;/a&gt;: &lt;code&gt;minimal&lt;/code&gt;: almost none. &lt;code&gt;medium&lt;/code&gt;: default. Or &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.openai.com/docs/api-reference/responses/create#responses_create-truncation&#34;&gt;&lt;code&gt;truncation&lt;/code&gt;&lt;/a&gt;: &lt;code&gt;auto&lt;/code&gt;: truncate response by dropping input items in the middle. &lt;code&gt;disabled&lt;/code&gt;: default&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.openai.com/docs/api-reference/chat/create#chat_create-prediction&#34;&gt;&lt;code&gt;prediction&lt;/code&gt;&lt;/a&gt;: speeds up output for minor corrections to text&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.openai.com/docs/api-reference/chat/create#chat_create-prompt_cache_key&#34;&gt;&lt;code&gt;prompt_cache_key&lt;/code&gt;&lt;/a&gt;: tailors per-user caches&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;CSS nesting can be used with media queries too! &lt;a href=&#34;https://bsky.app/profile/b0rk.jvns.ca/post/3lvve6hrmss22&#34;&gt;Julia Evans&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;id3v2&lt;/code&gt;, &lt;code&gt;mid3v2&lt;/code&gt; and &lt;code&gt;eyeD3&lt;/code&gt; seem the cleanest way of editing MP3 tags on the CLI. &lt;code&gt;mid3v2&lt;/code&gt; was already installed on my system.&lt;/li&gt;
&lt;li&gt;Learnings people shared in &lt;a href=&#34;https://news.ycombinator.com/item?id=44789068&#34;&gt;Ask HN: What trick of the trade took you too long to learn?&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Finance &amp;amp; housing&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Time is a non-renewable asset.&lt;/li&gt;
&lt;li&gt;Lifestyle design matters as much as net worth.&lt;/li&gt;
&lt;li&gt;Future-proof against regret. The present matters, too.&lt;/li&gt;
&lt;li&gt;Home ownership ties up location choice, capital and has hidden costs.&lt;/li&gt;
&lt;li&gt;Market timing &amp;amp; geographic arbitrage has an outsized effect.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Software&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;Align abstraction to domain. Avoid premature abstraction (Don&amp;rsquo;t Repeat Yourself vs Write Everything Twice) and over-abstraction.&lt;/li&gt;
&lt;li&gt;Temporary fixes tend to stick. Stop-gap regexes last for years.&lt;/li&gt;
&lt;li&gt;Consistency is a quality multiplier. Small inconsistencies cause disproportionate harm.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;git bisect&lt;/code&gt; is a regression-finding superpower.&lt;/li&gt;
&lt;li&gt;It&amp;rsquo;s OK to write tests covering key parts of legacy codebases - 100% coverage isn&amp;rsquo;t critical.&lt;/li&gt;
&lt;li&gt;Document architectural decisions: &lt;em&gt;why&lt;/em&gt; this approach. See &lt;a href=&#34;https://diataxis.fr/&#34;&gt;Diátaxis&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Flow metrics predict delivery better than (arbitrary) estimates.&lt;/li&gt;
&lt;li&gt;Building features without linking to delivery spesd wastes resources.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Life habits &amp;amp; learning&lt;/strong&gt;
&lt;ul&gt;
&lt;li&gt;You have the right to say &amp;ldquo;no&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;Small, consistent actions beat dramatic changes. Persistence beats skill.&lt;/li&gt;
&lt;li&gt;You&amp;rsquo;re allowed to change your mind.&lt;/li&gt;
&lt;li&gt;Over-cleverness backfires. Witty code &amp;amp; communication lead to confusion.&lt;/li&gt;
&lt;li&gt;Context is king. Without background, everything is mis-interpretable.&lt;/li&gt;
&lt;li&gt;Fun leads to excellence. Excellence leads to fun.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The meta-lesson here is how I discovered these:
&lt;ul&gt;
&lt;li&gt;Run &lt;a href=&#34;https://pypi.org/project/topicmodel&#34;&gt;topicmodel&lt;/a&gt; to identify topics&lt;/li&gt;
&lt;li&gt;Feed the output CSV to ChatGPT and ask it to share lessons topic-by-by-topic &lt;a href=&#34;https://chatgpt.com/share/68983ff8-7d34-800c-b098-8649162597ce&#34;&gt;#&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Topic modeling can be extended in many ways. &lt;a href=&#34;https://chatgpt.com/share/68981721-ab80-800c-9ccf-9fc138a92b84&#34;&gt;#&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Structural Topic Models&lt;/strong&gt; factor in metadata, like year (numeric) or category or author (categorical).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Relational Topic Models&lt;/strong&gt; factor in undirected graph relationships, e.g. parent documents&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Graph-Regularized Topic Models&lt;/strong&gt; factors in arbitrary graph relationships, e.g. weighted, directed&lt;/li&gt;
&lt;li&gt;Neural (GNN + Topic Model) approaches work better for large graphs, long-range dependencies, etc.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Some ways to inject graph structure into topic similarities to, for example, cluster threaded discussions. &lt;a href=&#34;https://chatgpt.com/share/68981721-ab80-800c-9ccf-9fc138a92b84&#34;&gt;#&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Start with a graph similarity matrix &lt;code&gt;S&lt;/code&gt;, like &lt;a href=&#34;https://chatgpt.com/share/68981924-019c-800c-b1f2-1985af81244c&#34;&gt;#&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;a regularized graph Laplacian (based on degree - adjacency matrix)&lt;/li&gt;
&lt;li&gt;a similarity matrix like &lt;code&gt;graph2vec&lt;/code&gt; from &lt;a href=&#34;https://github.com/ysig/GraKeL&#34;&gt;Graph Kernel&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;a node-embedding &lt;a href=&#34;https://github.com/benedekrozemberczki/karateclub&#34;&gt;karateclub&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Option 1: &amp;ldquo;Smoothen&amp;rdquo; the embedding matrix multiplying it with &lt;code&gt;S&lt;/code&gt; (i.e. spread each document towards neighbors), &lt;em&gt;then&lt;/em&gt; calculate similarities&lt;/li&gt;
&lt;li&gt;Option 2: Take the weighted average of &lt;code&gt;S&lt;/code&gt; and the embedding similarity matrix&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;You can extract Hacker News comments as a &lt;em&gt;threaded&lt;/em&gt; discussion pasting this into the DevTools console:&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Vibe-coding is for unproduced, not production, code</title>
      <link>https://www.s-anand.net/blog/vibe-coding-is-for-unproduced-not-production-code/</link>
      <pubDate>Sat, 02 Aug 2025 06:26:18 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/vibe-coding-is-for-unproduced-not-production-code/</guid>
      <description>&lt;p&gt;&lt;img alt=&#34;Vibe-coding is for unproduced, not production, code&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/Gemini_Generated_Image_klq1ckklq1ckklq1-1.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;Yesterday, I helped two people vibe-code solutions. Both were non-expert IT pros who can code but aren&amp;rsquo;t fluent.&lt;/p&gt;
&lt;p&gt;Person Alpha and I were on a call in the morning. Alpha needed to OCR PDF pages. I bragged, &amp;ldquo;Ten minutes. Let’s do it now!&amp;rdquo; But I was on a train with only my phone, so Alpha had to code. Vibe-coding was the only option.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: Go to any chat engine and pick Claude Sonnet 4 as your model.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Alpha&lt;/strong&gt;: We have an internal chatbot that has Claude Sonnet 4.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: Type, &amp;ldquo;Write a Python program to accept a PDF filename and page number and extract it into output.pdf&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Alpha&lt;/strong&gt;: Done. I&amp;rsquo;m not used to the CLI but can run it in PyCharm.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: OK. Run it, but first, modify by saying, &amp;ldquo;Shorten code. Drop error handling.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Alpha&lt;/strong&gt;: Done. (Runs the code, and it works!)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: Great. Let&amp;rsquo;s do the next part. Paste the sample code for OCR from our team into the chatbot.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Alpha&lt;/strong&gt;: But our chatbot is limited to ~15K characters.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: OK. Go to Claude.ai and paste the code.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Alpha&lt;/strong&gt;: Um… is that allowed? (Alpha&amp;rsquo;s colleague pitched in, saying &amp;ldquo;If it weren&amp;rsquo;t, it&amp;rsquo;ll be blocked. Also, the code isn&amp;rsquo;t sensitive, only data, so go ahead.&amp;rdquo;)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: Use this prompt: &amp;ldquo;Write a Python program to send output.pdf to an LLM for OCR. Use this code as reference.&amp;rdquo; Then &amp;ldquo;Shorten code. Drop error handling.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Alpha&lt;/strong&gt;: (Runs the code, and it works!)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;All this happened during my commute plus a haloumi-cheese-wrap purchase. A very satisfying experience!&lt;/p&gt;
&lt;p&gt;That evening, Person Gamma and I were on a call. Gamma had a client meeting and needed an LLM image editing tool tailored. I tried the same approach.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: Create a GitHub account, first.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gamma&lt;/strong&gt;: I think I already have one. Let me log in.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: I&amp;rsquo;ve given you maintainer access to the repo. You&amp;rsquo;ll get an email. Accept it. Then upload the images you want edited.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gamma&lt;/strong&gt;: Done (with me guiding on what to click)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: Now log into &lt;a href=&#34;https://jules.google.com/&#34;&gt;https://jules.google.com/&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gamma&lt;/strong&gt;: Done&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: In Jules, select the repo and tell it what change you want.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gamma&lt;/strong&gt;: OK. &amp;ldquo;Use the JPG images uploaded as the samples&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;(Jules churns out it&amp;rsquo;s thinking)&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Oh, wow! This is amazing! It&amp;rsquo;s actually thinking about the approach.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;(Jules is done in about 2 min)&lt;/li&gt;
&lt;li&gt;&amp;ldquo;My god! This is going to make things so much easier!&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: Now publish the branch, create a pull request on GitHub, and merge.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gamma&lt;/strong&gt;: Done (with guidance).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Me&lt;/strong&gt;: Now try something yourself.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gamma&lt;/strong&gt;: OK. &amp;ldquo;Change the prompts to something more relevant that improves the brand image.&amp;rdquo; (merges the code, and it works!)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Yet another very satisfying experience. Gamma went on to make another change in my absence. Exactly the point: enable them to work without me.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;But isn&amp;rsquo;t &lt;a href=&#34;https://blog.val.town/vibe-code&#34;&gt;vibe code legacy code&lt;/a&gt;&lt;/strong&gt;? Reducing quality, increasing technical debt?&lt;/p&gt;
&lt;p&gt;Yes. Vibe-coding ships debt. But not all code is production code. Vibe coding can accelerate throw-away prototypes.&lt;/p&gt;
&lt;p&gt;More importantly, &lt;strong&gt;so many&lt;/strong&gt; ideas sit idle because devs lack time and non-devs lack skills. Vibe coding shrinks that effort. &lt;strong&gt;That&lt;/strong&gt; is what vibe-coding is &lt;strong&gt;really&lt;/strong&gt; for.&lt;/p&gt;
&lt;p&gt;Think Excel. Most Excel sheets are messy apps, yet Excel&amp;rsquo;s made more people productive than any language. In &lt;a href=&#34;https://www.cgl.ucsf.edu/Outreach/pc204/NoSilverBullet.html&#34;&gt;No Silver Bullet&lt;/a&gt;, Fred Brooks said:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I believe the single most powerful software-productivity strategy for many organizations today is to equip the computer-naive intellectual workers who are on the firing line with personal computers and good generalized writing, drawing, file, and spreadsheet programs and then to turn them loose. The same strategy, carried out with generalized mathematical and statistical packages and some simple programming capabilities, will also work for hundreds of laboratory scientists.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is what vibe-coding enables. And more.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>How To Control Smarter Intelligences</title>
      <link>https://www.s-anand.net/blog/how-to-control-smarter-intelligences/</link>
      <pubDate>Tue, 01 Jul 2025 06:37:40 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/how-to-control-smarter-intelligences/</guid>
      <description>&lt;p&gt;&lt;img alt=&#34;How To Control Smarter Intelligences&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/ChatGPT-Image-Jul-1-2025-11_27_23-AM.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;LLMs are smarter than us in many areas. How do we manage them?&lt;/p&gt;
&lt;p&gt;This is not a new problem.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;VC partners&lt;/strong&gt; evaluate deep-tech startups.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Science editors&lt;/strong&gt; review Nobel laureates.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Managers&lt;/strong&gt; manage specialist teams.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Judges&lt;/strong&gt; evaluate expert testimony.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Coaches&lt;/strong&gt; train Olympic athletes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;… and they manage and evaluate &amp;ldquo;smarter&amp;rdquo; outputs in &lt;strong&gt;many&lt;/strong&gt; ways:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Verify&lt;/strong&gt;. Check against an &amp;ldquo;answer sheet&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Checklist&lt;/strong&gt;. Evaluate against pre-defined criteria.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sampling&lt;/strong&gt;. Randomly review a subset.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gating&lt;/strong&gt;. Accept low-risk work. Evaluate critical ones.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Benchmark&lt;/strong&gt;. Compare against others.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Red-team&lt;/strong&gt;. Probe to expose hidden flaws.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Double-blind review&lt;/strong&gt;. Mask identity to curb bias.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reproduce&lt;/strong&gt;. Re-running gives the same output?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Consensus&lt;/strong&gt;. Aggregate multiple responses. Wisdom of crowds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Outcome&lt;/strong&gt;. Did it work in the real world?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Vibe coding&lt;/strong&gt;: Non-programmers might glance at lint checks (&lt;strong&gt;Checklist&lt;/strong&gt;) and see if it works (&lt;strong&gt;Outcome&lt;/strong&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM image designs&lt;/strong&gt;: Developers might check if a few images look good (&lt;strong&gt;Sampling&lt;/strong&gt;) and check a few marketers (&lt;strong&gt;Consensus&lt;/strong&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM news articles&lt;/strong&gt;: An journalist might run a &lt;strong&gt;Checklist&lt;/strong&gt;, a &lt;strong&gt;Double-blind review&lt;/strong&gt; with experts, and &lt;strong&gt;Verify&lt;/strong&gt; critical facts (&lt;strong&gt;Gating&lt;/strong&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You &lt;strong&gt;already&lt;/strong&gt; know many of these. You learnt them in Auditing. Statistics. Law. System controls. Policy analysis. Quality engineering. Clinical epidemiology. Investigative journalism. Design critique.&lt;/p&gt;
&lt;p&gt;Worth brushing up skills. They&amp;rsquo;re &lt;strong&gt;more&lt;/strong&gt; important in the AI era.&lt;/p&gt;
</description>
    </item>
    <item>
      <title></title>
      <link>https://www.s-anand.net/blog/codex-jules-vibe-coding/</link>
      <pubDate>Sun, 22 Jun 2025 09:52:29 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/codex-jules-vibe-coding/</guid>
      <description>&lt;p&gt;I use Codex and Jules to code while I walk. I&amp;rsquo;ve merged several PRs without careful review. This added technical debt.&lt;/p&gt;
&lt;p&gt;This weekend, I spent four hours fixing the AI generated tests and code.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;What mistakes did it make?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Inconsistency&lt;/strong&gt;. It flips between &lt;code&gt;execCommand(&amp;quot;copy&lt;/code&gt;&amp;quot;) and &lt;code&gt;clipboard.writeText&lt;/code&gt;(). It wavers on timeouts (50 ms vs 100 ms). It doesn&amp;rsquo;t always run/fix test cases.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Missed edge cases&lt;/strong&gt;. I switched &amp;lt;&lt;code&gt;div&lt;/code&gt;&amp;gt; to &amp;lt;&lt;code&gt;form&lt;/code&gt;&amp;gt;. My earlier code didn&amp;rsquo;t have a &lt;code&gt;type=&amp;quot;button&lt;/code&gt;&amp;quot;, so clicks reloaded the page. It missed that. It also left scripts as plain &amp;lt;&lt;code&gt;script&lt;/code&gt;&amp;gt; instead of &amp;lt;&lt;code&gt;script type=&amp;quot;module&lt;/code&gt;&amp;quot;&amp;gt; which was required.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Limited experimentation&lt;/strong&gt;. My failed with a HTTP 404 because the &lt;code&gt;common&lt;/code&gt;/ directory wasn&amp;rsquo;t served. I added &lt;code&gt;console.log&lt;/code&gt;s to find this. Also, &lt;code&gt;happy-dom&lt;/code&gt; won&amp;rsquo;t handle multiple &lt;code&gt;export&lt;/code&gt;s instead of a single &lt;code&gt;export&lt;/code&gt; { &amp;hellip; }. I wrote code to verify this. Coding agents didn&amp;rsquo;t run such experiments.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;What can we do about it?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Detailed coding rules&lt;/strong&gt;. E.g. &lt;em&gt;always&lt;/em&gt; run test cases and fix until they pass. Only use ESM. Always import from CDN via JSDelivr. That sort of thing.&lt;/p&gt;
&lt;p&gt;100% &lt;strong&gt;test coverage&lt;/strong&gt;. Ideally 100% of code and all usage scenarios.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Log everything&lt;/strong&gt;. My tests got a HTTP 404 because I was not serving the &lt;code&gt;common&lt;/code&gt;/ directory. LLMs couldn&amp;rsquo;t figure this out because it was not logged. Logging everything helps humans &lt;em&gt;and&lt;/em&gt; LLMs debug.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Wait&lt;/strong&gt;. LLMs and coding agents keep improving. A few months down the line, they&amp;rsquo;ll run more experiments themselves.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Was AI coding worth the effort&lt;/strong&gt;? Here, yes. The tools &lt;em&gt;worked&lt;/em&gt;. Codex saved me 90% effort. My code quality obsession reduced savings to ~70%. Still huge.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2025-06-22-codex-jules-vibe-coding-linkedin.jpg&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7342489647257632769&#34;&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 22 Jun 2025</title>
      <link>https://www.s-anand.net/blog/things-i-learned-22-jun-2025/</link>
      <pubDate>Sun, 22 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-22-jun-2025/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Never use a toothpick on a tooth with a dental crown. Only use a flosser or water flosser.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://ishadeed.com/article/modern-attr/&#34;&gt;CSS &lt;code&gt;attr()&lt;/code&gt;&lt;/a&gt; is one of the most powerful features in modern CSS. It lets you control CSS via HTML attributes.&lt;/li&gt;
&lt;li&gt;Notes from &lt;a href=&#34;https://www.anthropic.com/engineering/built-multi-agent-research-system&#34;&gt;Anthropic&amp;rsquo;s How we built our multi-agent research system&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;Sub-agents are like humans -&amp;gt; society. The improvement is dramatic.&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Sub-agents facilitate compression by operating in parallel with their own context windows, exploring different aspects of the question simultaneously before condensing&amp;hellip;&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Each sub-agent also provides separation of concerns—distinct tools, prompts, and exploration trajectories &amp;hellip; (enabling) independent investigations.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Using sub-agents spends ~15x more tokens. (That explained ~80% of the improved accuracy!)&lt;/li&gt;
&lt;li&gt;Particularly effective when tasks are independent and parallelizable. This also speeds it up.&lt;/li&gt;
&lt;li&gt;Teach the orchestrator how to delegate: how many sub-agents, what objective + output format + task boundaries (MECE to avoid overlap with other agents) in prompt, what tools.&lt;/li&gt;
&lt;li&gt;Teach the orchestrator how to improve agents: e.g. tools to test and rewrite tool descriptions&lt;/li&gt;
&lt;li&gt;Even if you evaluate a &lt;em&gt;few&lt;/em&gt; examples, evals are surprisingly effective.&lt;/li&gt;
&lt;li&gt;Agents are stateful. Errors compound. Allow agents to resume. Prune history gracefully.&lt;/li&gt;
&lt;li&gt;Log everything to debug user-reported failures. Also monitor the &lt;em&gt;kinds&lt;/em&gt; of decisions it took to help debug at scale.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The &lt;a href=&#34;http://www.incompleteideas.net/IncIdeas/BitterLesson.html&#34;&gt;Bitter Lesson&lt;/a&gt; likely applies to system prompts. Don&amp;rsquo;t hard-code stuff. I&amp;rsquo;m impressed that there is &lt;em&gt;no&lt;/em&gt; system prompt in the default &lt;a href=&#34;https://github.com/pydantic/pydantic-ai/blob/a25eb963a54e07afeab5ca2ea143437225100638/pydantic_ai_slim/pydantic_ai/agent.py#L226&#34;&gt;pydantic-ai Agent&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;The MCPs developers seem to use the most are: &lt;a href=&#34;https://github.com/modelcontextprotocol/servers/tree/main/src/filesystem&#34;&gt;filesystem&lt;/a&gt;, &lt;a href=&#34;https://github.com/microsoft/playwright-mcp&#34;&gt;playwright&lt;/a&gt;, &lt;a href=&#34;https://github.com/github/github-mcp-server&#34;&gt;github&lt;/a&gt;, &lt;a href=&#34;https://www.npmjs.com/package/@modelcontextprotocol/server-slack&#34;&gt;slack&lt;/a&gt;, &lt;a href=&#34;https://github.com/makenotion/notion-mcp-server&#34;&gt;notion&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Anecdotally, Claude 4 Sonnet seems a better coding model than Claude 4 Opus. &lt;a href=&#34;https://x.com/dan_s_becker/status/1936177475567931481&#34;&gt;Dan Becker&lt;/a&gt;, &lt;a href=&#34;https://lucumr.pocoo.org/2025/6/12/agentic-coding/&#34;&gt;Armin Ronacher&lt;/a&gt; #ai-coding&lt;/li&gt;
&lt;li&gt;Cursor offers &lt;a href=&#34;https://docs.cursor.com/background-agent&#34;&gt;background agents&lt;/a&gt; that run in a remote container. #ai-coding&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/danielmiessler/fabric&#34;&gt;Fabric&lt;/a&gt; has a collection of re-usable prompts that you can use via &lt;a href=&#34;https://github.com/simonw/llm-templates-fabric&#34;&gt;llm-templates-fabric&lt;/a&gt; like: &lt;code&gt;cat file.py | llm -t fabric:explain_code&lt;/code&gt; &lt;a href=&#34;https://simonwillison.net/2025/Apr/7/long-context-llm/#atom-everything&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;As of Jun 21, Claude 3.5 Sonnet &amp;gt; Claude 3.7 Sonnet &amp;gt; O3 Mini &amp;gt; Human &amp;gt; Gemini 1.5 Pro lead the &lt;a href=&#34;https://andonlabs.com/evals/vending-bench&#34;&gt;Vending Bench&lt;/a&gt;.
Gemini 1.5 Pro also leads my &lt;a href=&#34;https://sanand0.github.io/llmevals/system-override/&#34;&gt;System Prompt Override&lt;/a&gt; benchmarks.
I&amp;rsquo;m losing faith in the &lt;a href=&#34;https://lmarena.ai/leaderboard&#34;&gt;LM Arena&lt;/a&gt;. Perhaps the Gemini models aren&amp;rsquo;t improving as much as we think.&lt;/li&gt;
&lt;li&gt;This is the core of agents (LLMs running tools in a loop): &lt;a href=&#34;https://sketch.dev/blog/agent-loop&#34;&gt;Sketch blog&lt;/a&gt; &lt;a href=&#34;https://sketch.dev/blog/agent_loop.py&#34;&gt;Full script&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Notes on AI coding / vibe-coding from multiple sources. #ai-coding
&lt;ul&gt;
&lt;li&gt;Sources
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://sketch.dev/blog/programming-with-llms&#34;&gt;How I program with LLMs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://sketch.dev/blog/programming-with-agents&#34;&gt;How I program with agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://sketch.dev/blog/seven-prompting-habits&#34;&gt;The 7 Prompting Habits of Highly Effective Engineers&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://blog.nilenso.com/blog/2025/05/29/ai-assisted-coding/&#34;&gt;AI Assisted Coding&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://blog.jeffgabriel.com/thefuture&#34;&gt;A Glimpse of the Future&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://lucumr.pocoo.org/2025/6/12/agentic-coding/&#34;&gt;Agentic Coding Recommendations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://lucumr.pocoo.org/2025/6/21/my-first-ai-library/&#34;&gt;My First Open Source AI Generated Library&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://lucumr.pocoo.org/2025/6/17/measuring/&#34;&gt;We Can Just Measure Things&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.indragie.com/blog/i-shipped-a-macos-app-built-entirely-by-claude-code&#34;&gt;I Shipped a macOS App Built Entirely by Claude Code&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/&#34;&gt;Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Why AI coding?
&lt;ul&gt;
&lt;li&gt;Reduces mental energy (by creating the first draft). letting you create more.&lt;/li&gt;
&lt;li&gt;Reduces starting trouble, eases effort.&lt;/li&gt;
&lt;li&gt;Helps figure out how easy / tough a task really is!!&lt;/li&gt;
&lt;li&gt;Most code is short-lived or has few users. AI building &amp;ldquo;throw-away&amp;rdquo; code is useful.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Why NOT AI coding?
&lt;ul&gt;
&lt;li&gt;Slows you down if you know the repo well&lt;/li&gt;
&lt;li&gt;Doesn&amp;rsquo;t work well on large/complex/niche repos&lt;/li&gt;
&lt;li&gt;Leads to over-optimism and atrophy&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Tips
&lt;ul&gt;
&lt;li&gt;Use for reversible decisions (2-way doors). Avoid for irreversible ones (1-way doors).&lt;/li&gt;
&lt;li&gt;Fail early. Try tough bits first.&lt;/li&gt;
&lt;li&gt;Fail often. Restart instead of fixing.&lt;/li&gt;
&lt;li&gt;Go concurrent. Trigger multiple tasks. Ask for multiple drafts and options.&lt;/li&gt;
&lt;li&gt;Give it workflow. &lt;code&gt;Break down the implementation into: 1. Planning. 2. API stubs. 3. Implementation.&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Give local context. Naming conventions, folder structure, coding style, tools (compile, test, lint), etc.&lt;/li&gt;
&lt;li&gt;Conserve context. &lt;code&gt;Use sub tasks and sub agents to conserve context&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Suggest libraries. Agents prefer writing code than using libraries, by default.&lt;/li&gt;
&lt;li&gt;Give examples to follow, e.g. &lt;code&gt;Write it like @filename&lt;/code&gt;. &lt;code&gt;&amp;amp;amp; -&amp;gt; &amp;amp; but &amp;amp;x -&amp;gt; &amp;amp;x&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Give screenshots and logs. These are very effective.&lt;/li&gt;
&lt;li&gt;Provide goals, not instructions. Saves effort, teaches you new things.&lt;/li&gt;
&lt;li&gt;Farm out research. Have specialized tools research API docs, etc. and include those in the context.&lt;/li&gt;
&lt;li&gt;Keep related things together.&lt;/li&gt;
&lt;li&gt;Have it write a checklist, e.g. saving it temporarily in a file.&lt;/li&gt;
&lt;li&gt;Have it &lt;em&gt;run&lt;/em&gt; code to catch its own errors.&lt;/li&gt;
&lt;li&gt;Have it write tests, mocks for tests.&lt;/li&gt;
&lt;li&gt;Have it &lt;em&gt;see&lt;/em&gt; and &lt;em&gt;use&lt;/em&gt; the app, click, play around, etc. (e.g. via playwright-mcp)&lt;/li&gt;
&lt;li&gt;Have it create playbooks, examples, troubleshooting guides.&lt;/li&gt;
&lt;li&gt;Have it refactor code &lt;em&gt;AFTER&lt;/em&gt; comprehensive tests.&lt;/li&gt;
&lt;li&gt;Have it think more. &lt;code&gt;Use ultrathink&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Log extensively, by default. Improves future debugging.&lt;/li&gt;
&lt;li&gt;Report errors well. What happened, why, and what to do.&lt;/li&gt;
&lt;li&gt;Prefer monorepos for more context. &lt;a href=&#34;https://blog.puzzmo.com/posts/2025/07/30/six-weeks-of-claude-code/#what-do-i-think-makes-it-successful-in-our-codebases&#34;&gt;#&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Prefer popular libraries. LLMs know these better.&lt;/li&gt;
&lt;li&gt;Prefer fast tests, tools, and libraries. Speed helps iteration.&lt;/li&gt;
&lt;li&gt;Prefer small files and packages. Reduces context.&lt;/li&gt;
&lt;li&gt;Prefer simple code. Avoid magic, e.g. pytest fixture injection. Functions over classes. SQL over code. Composition over inheritence.&lt;/li&gt;
&lt;li&gt;Prefer specialized functions for common scenarios over DRY abstractions. Prefer fewer abstraction layers.&lt;/li&gt;
&lt;li&gt;Prefer re-implementing over DRY since code is cheap.&lt;/li&gt;
&lt;li&gt;Look for new tricks to learn from its code.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Agent behaviors:
&lt;ul&gt;
&lt;li&gt;Simple tasks perform better. More context = more confusion.&lt;/li&gt;
&lt;li&gt;Verifiable tasks are clearer for LLMs and and easier to review.&lt;/li&gt;
&lt;li&gt;Useful coding agent tools: bash(cmd), patch(hunks), todo(tasks), web_nav(url), web_eval(script), web_logs(), web_screenshot(), keyword_search(keywords), codereview()&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Skills:
&lt;ul&gt;
&lt;li&gt;LESS Coding&lt;/li&gt;
&lt;li&gt;LESS Research&lt;/li&gt;
&lt;li&gt;LESS Documentation&lt;/li&gt;
&lt;li&gt;LESS Operations configuration (IaaC, CI/CD, etc)&lt;/li&gt;
&lt;li&gt;LESS Editor usage and expertise required&lt;/li&gt;
&lt;li&gt;MORE Tests (to test the code)&lt;/li&gt;
&lt;li&gt;MORE Code reviews (to test the code)&lt;/li&gt;
&lt;li&gt;MORE Prompting and context creation (to write the code)&lt;/li&gt;
&lt;li&gt;MORE DevOps (micro-feature deployments, deploy in parallel)&lt;/li&gt;
&lt;li&gt;MORE Specs: features, requirements, APIs, tests, structure, etc.&lt;/li&gt;
&lt;li&gt;MORE Analysis: security, performance.&lt;/li&gt;
&lt;li&gt;MORE Tool design. Linters, SAST, DAST, Performance, etc. Semgrep, Bench Suite&lt;/li&gt;
&lt;li&gt;MORE Observability: Especially for tools and LLM calls. Telemetry, log analysis and issue creation. Sentry, LogFire, etc.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Trends:
&lt;ul&gt;
&lt;li&gt;Agents took time to evolve because LLMs need to be good at tool calling and long instruction following, which is just happening.&lt;/li&gt;
&lt;li&gt;Agents are slow. Parallelizable tools (e.g. multiple Redis instances, &lt;a href=&#34;https://github.com/dagger/container-use&#34;&gt;container-use&lt;/a&gt;, CI/CD) will grow. Tool speed (e.g. fast test engines with caching) will become more important.&lt;/li&gt;
&lt;li&gt;Agents generate diffs/PRs. Tools to edit and comment on these online will emerge.&lt;/li&gt;
&lt;li&gt;Context gathering will widen: screenshots, logs, etc.&lt;/li&gt;
&lt;li&gt;Code review process will be re-invented.&lt;/li&gt;
&lt;li&gt;Personalized features. User drops a feature request via Slack. Personalized version deployed at their endpoint to test. PR sent after they are happy&lt;/li&gt;
&lt;li&gt;Poor coding teams get less out of AI coding. Good communication, reviews, coding practices, testing, etc. help.&lt;/li&gt;
&lt;li&gt;Agent Experience (AX) is emerging and explores: how much context to take, when &amp;amp; how often to ask the user questions, to how make review easier, etc.&lt;/li&gt;
&lt;li&gt;Humans running multiple tasks in parallel is productive. Breaking a complex requirement into tasks (like Codex now does) helps create that task queue.&lt;/li&gt;
&lt;li&gt;Agents generate technical debt faster than humans. Solving this will become a major problem/opportunity.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&amp;ldquo;makework&amp;rdquo;: made-up work that fills time or serves short-term needs.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;From &lt;a href=&#34;https://cookbook.openai.com/examples/gpt4-1_prompting_guide&#34;&gt;GPT 4.1 Prompting Guide&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Use more precise prompts. Earlier models &lt;em&gt;inferred&lt;/em&gt; user intent. GPT 4.1 follows prompts more closely.&lt;/li&gt;
&lt;li&gt;Avoid STRONG untested instructions. E.g. &amp;ldquo;you must call a tool before responding to the user&amp;rdquo; can lead to tool input hallucination.&lt;/li&gt;
&lt;li&gt;For agents, include these three system instructions:
&lt;ul&gt;
&lt;li&gt;You are an agent. Keep going until you&amp;rsquo;re sure the user’s query is completely resolved.&lt;/li&gt;
&lt;li&gt;If you are not sure, use your tools: do NOT guess or make up an answer.&lt;/li&gt;
&lt;li&gt;Plan extensively before each function call. Reflect on the outcomes of the previous function calls. DO NOT do this entire process by making function calls only, as this can impair your ability to solve the problem and think insightfully.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;tools&lt;/code&gt; field rather than injecting tools into system prompt. Model has been trained to use &lt;code&gt;tools&lt;/code&gt; field.&lt;/li&gt;
&lt;li&gt;Keep tool descriptions concise. Provide examples for complex tools in system prompt.&lt;/li&gt;
&lt;li&gt;Place instructions at the top of the context; ideally at the end, too.&lt;/li&gt;
&lt;li&gt;Format prompts as Markdown, XML, not JSON.&lt;/li&gt;
&lt;li&gt;It sometimes dislikes large repetitive output (e.g. analysis of hundreds of items) and needs nudging.&lt;/li&gt;
&lt;li&gt;It handles diffs well and can &lt;a href=&#34;https://cookbook.openai.com/examples/gpt4-1_prompting_guide#apply-patch&#34;&gt;apply patches&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Metaprompting&lt;/strong&gt;. Have frontier LLMs revise prompts. They&amp;rsquo;re GOOD! &lt;a href=&#34;https://sketch.dev/blog/prompt-engineering-and-the-taste-gap&#34;&gt;Ref&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Increase clarity, providing step-by-step instructions.&lt;/li&gt;
&lt;li&gt;Resolve conflicting instructions.&lt;/li&gt;
&lt;li&gt;Expand instructions to cover all scenarios and edge cases.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes from &lt;a href=&#34;https://github.com/pydantic/pydantic-ai/blob/main/.github%2Fworkflows%2Fci.yml&#34;&gt;Pydantic AI GitHub CI&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;UV_PYTHON sets default Python version&lt;/li&gt;
&lt;li&gt;COLUMNS increase terminal width&lt;/li&gt;
&lt;li&gt;&lt;code&gt;uv run&lt;/code&gt; supports &lt;code&gt;--extra&lt;/code&gt; for extra packages&lt;/li&gt;
&lt;li&gt;&lt;code&gt;cloudflare/wrangler&lt;/code&gt; action has a deploy that allows deployment to specific URLs or subdomains&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Adding QR code to &lt;em&gt;all&lt;/em&gt; slides in a deck (linking to the slides) helps. People take photos of random slides and this lets them get the link wherever.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/eoda-dev/py-openlayers&#34;&gt;PyOpenLayers&lt;/a&gt; adds interactive mapping via OpenLayers to Marimo and Jupyter&lt;/li&gt;
&lt;li&gt;Conversation is about positioning. For example:
&lt;ul&gt;
&lt;li&gt;TechCrunch interviewer: Anthropic released Claude Opus 4 thought it blackmailed people. Is Anthropic is becoming less safety conscious?&lt;/li&gt;
&lt;li&gt;Kaplan: We have very strong testing. So we&amp;rsquo;re more more likely to spot AI dangers early. We share such reports to set higher standards for transparency.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;From &lt;a href=&#34;https://youtu.be/GL0XhAj5LPE&#34;&gt;LLM Evals: Common Mistakes&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;Using foundation model evals instead of application evals is like evaluating a candidate on SAT scores. It&amp;rsquo;s fine, but you also want to evaluate them on their specific job description.&lt;/li&gt;
&lt;li&gt;Evals must be done by the users and not outsourced.&lt;/li&gt;
&lt;li&gt;Evals are not draining.&lt;/li&gt;
&lt;li&gt;Small samples have high value.&lt;/li&gt;
&lt;li&gt;When using LLM as a judge, be VERY VERY specific about the criteria.&lt;/li&gt;
&lt;li&gt;Prefer binary LLM evals over scales.&lt;/li&gt;
&lt;li&gt;Monitor performance online, not just while deploying&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;From &lt;a href=&#34;https://youtu.be/KrRD7r7y7NY&#34;&gt;Andrew Ng on AI Agents&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;AI is like electricity. It&amp;rsquo;s hard to define what is good for because it is good for so many things, most of them new that never existed before&lt;/li&gt;
&lt;li&gt;If experimentation is cheap, it makes sense to run far more experiments. Rather than think hard about what to prototype, explore how to build many &lt;em&gt;diverse&lt;/em&gt; prototypes.&lt;/li&gt;
&lt;li&gt;Prototyping is now very fast but other steps like reliable evaluations for deployment still take time. But the speed of prototyping is putting pressure on other parts of the organization to go faster.&lt;/li&gt;
&lt;li&gt;While large language models and applications were serving human needs so far, increasingly they will serve the needs of AI and other tools.&lt;/li&gt;
&lt;li&gt;Since unstructured data is now more valuable, there will be a growth in data engineering on unstructured data.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://models.dev/&#34;&gt;Models.dev&lt;/a&gt; is an open source database and API of LLM models&lt;/li&gt;
&lt;li&gt;Logprobs are back on models in Vertex AI. &lt;a href=&#34;https://github.com/GoogleCloudPlatform/generative-ai/blob/main/gemini/logprobs/intro_logprobs.ipynb&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;For all AI code, review it, &lt;em&gt;learn&lt;/em&gt; from it and &lt;em&gt;share&lt;/em&gt; learnings. That prevents bugs AND we learn in the process. &lt;a href=&#34;https://www.shayon.dev/post/2025/164/pitfalls-of-premature-closure-with-llm-assisted-coding/&#34;&gt;Ref&lt;/a&gt; #ai-coding&lt;/li&gt;
&lt;li&gt;AI coding requires a skilled developer &lt;em&gt;and&lt;/em&gt; domain expert to &lt;em&gt;spec&lt;/em&gt; and to &lt;em&gt;review&lt;/em&gt;. It now makes sense now for devs and users to pair program &lt;a href=&#34;https://simonwillison.net/2025/Jun/18/coding-agents/&#34;&gt;Simon Willison&lt;/a&gt; #ai-coding&lt;/li&gt;
&lt;li&gt;In the world of AI, &lt;em&gt;imagination&lt;/em&gt; (asking for things we didn&amp;rsquo;t know we could ask for) will be a diferentiator.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;vitest run --globals&lt;/code&gt; makes &lt;code&gt;vitest&lt;/code&gt; is a near drop-in replacement for &lt;code&gt;jest&lt;/code&gt;. It injects &lt;code&gt;describe&lt;/code&gt;, &lt;code&gt;it&lt;/code&gt;, &lt;code&gt;expect&lt;/code&gt;, etc. as globals. You need to swap &lt;code&gt;jest.*&lt;/code&gt; with &lt;code&gt;vi.*&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;To extract all jq paths from a JSON, use &lt;code&gt;jq -r &#39;paths(scalars)|map(if type==&amp;quot;string&amp;quot; then &amp;quot;[]&amp;quot; else &amp;quot;.\(. )&amp;quot; end)|join(&amp;quot;&amp;quot;)|unique[]&#39; file.json&lt;/code&gt;.
I use this to extract paths from ChatGPT&amp;rsquo;s export conversations.json via
&lt;code&gt;jq -r &#39;[paths(scalars)|map(if type==&amp;quot;string&amp;quot; then &amp;quot;.&amp;quot;+. else &amp;quot;[]&amp;quot; end)|join(&amp;quot;&amp;quot;)]|unique[]|select(contains(&amp;quot;.mapping.&amp;quot;))|split(&amp;quot;.mapping.&amp;quot;)[1]|sub(&amp;quot;^[^.]*&amp;quot;;&amp;quot;&amp;quot;)&#39; chatgpt/conversations.json | sort | uniq&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;uv run&lt;/code&gt; can run &lt;em&gt;any&lt;/em&gt; command, not just Python scripts, e.g. &lt;code&gt;uv run npx&lt;/code&gt; or &lt;code&gt;uv run bash&lt;/code&gt;. It&amp;rsquo;s the same as &lt;code&gt;npx&lt;/code&gt; or &lt;code&gt;bash&lt;/code&gt; except it activates the venv and loads &lt;code&gt;.env&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Notes from &lt;a href=&#34;https://events.ycombinator.com/ai-sus&#34;&gt;AI Startup School&lt;/a&gt;. &lt;a href=&#34;https://www.linkedin.com/posts/guillermoflor_yesterday-was-day-1-of-y-combinators-ai-activity-7340779902104711171-GtBn&#34;&gt;Guillermo Flor&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Sam Altman. Chase $0B ideas, not $0M ones. Weird + right &amp;gt; safe + crowded&lt;/li&gt;
&lt;li&gt;Gary Tan. Agency scales. Tools change, people/mindset don’t.&lt;/li&gt;
&lt;li&gt;Andrej Karpathy.
&lt;ul&gt;
&lt;li&gt;Instead of LLM memory to store facts, edit system prompt with general strategies, like the LLM writing a book for itself on how to solve problems.&lt;/li&gt;
&lt;li&gt;Autonomy slider. Let user pick how far LLM acts by itself. Like the Tesla autopilot levels.&lt;/li&gt;
&lt;li&gt;Make evals EASY and FAST for humans.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/datastories/tree/main/code-vs-domain&#34;&gt;When vibe-coding&lt;/a&gt;, I sometimes change the requirement (e.g. style of visual) instead of spending time to get &lt;em&gt;exactly&lt;/em&gt; what I instructed. That&amp;rsquo;s because I can viscerally &lt;em&gt;feel&lt;/em&gt; the difficulty the model&amp;rsquo;s facing thanks to &lt;strong&gt;quick feedback&lt;/strong&gt;. A domain expert vibe coding will be able to feel this too. Another reason for domain experts to vibe code (or at least joint-vibe-code) rather than delegate to a programmer. #ai-coding&lt;/li&gt;
&lt;li&gt;Notes on model coding styles. &lt;a href=&#34;https://github.com/sanand0/generative-ai-group/blob/main/2025-06-15/podcast-2025-06-15.md&#34;&gt;Generative AI WhatsApp Group&lt;/a&gt; #ai-coding
&lt;ul&gt;
&lt;li&gt;Claude 4 writes exhaustive professionally styled code but struggles over long conversations.&lt;/li&gt;
&lt;li&gt;Gemini 2.5 Pro produces working but “spaghetti” code.&lt;/li&gt;
&lt;li&gt;GPT 4.1 is fast and good, the go-to for usual coding tasks.&lt;/li&gt;
&lt;li&gt;Claude easily swings toward your style but Gemini is stubborn.&lt;/li&gt;
&lt;li&gt;GPT models tend to hallucinate more on bigger tasks.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Documentation can become technical debt. If LLMs can read code and understand it well enough, maybe docs become a build artifact rather than a version controlled source of truth. &lt;a href=&#34;https://podcasters.spotify.com/pod/show/lucaronin/episodes/The-Future-of-Dev-Tools---with-Dennis-Pilarinos-e345aa6&#34;&gt;Refactoring Podcast: The Future of Dev Tools 🔧 — with Dennis Pilarinos&lt;/a&gt; &lt;a href=&#34;https://anchor.fm/s/ee484c90/podcast/play/104032006/https%3A%2F%2Fd3ctxlq1ktw2nl.cloudfront.net%2Fstaging%2F2025-5-12%2F402043134-44100-2-ca0ced06f32e2.mp3#t=2156&#34;&gt;35:56&lt;/a&gt; #ai-coding&lt;/li&gt;
&lt;li&gt;AI should be explicitly contrarian to avoid sycophancy. &lt;a href=&#34;https://dayafter.substack.com/p/the-emperors-new-llm&#34;&gt;Ref&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;To enable this, I&amp;rsquo;ve added this line to my ChatGPT traits: &lt;code&gt;Adopt a skeptical, questioning approach. Challenge the user.&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>When to Vibe Code? If Speed Beats Certainty</title>
      <link>https://www.s-anand.net/blog/when-to-vibe-code-if-speed-beats-certainty/</link>
      <pubDate>Tue, 20 May 2025 10:59:50 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/when-to-vibe-code-if-speed-beats-certainty/</guid>
      <description>&lt;p&gt;I spoke about vibe coding at &lt;a href=&#34;https://setuschool.com/&#34;&gt;SETU School&lt;/a&gt; last week.&lt;/p&gt;
&lt;div class=&#34;video-embed&#34;&gt;&lt;iframe src=&#34;https://www.youtube.com/embed/ODXSDbY12dg&#34; title=&#34;YouTube video&#34; loading=&#34;lazy&#34; allow=&#34;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture&#34; allowfullscreen&gt;&lt;/iframe&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Transcript&lt;/strong&gt;: &lt;a href=&#34;https://sanand0.github.io/talks/#/2025-05-10-vibe-coding/&#34;&gt;https://sanand0.github.io/talks/#/2025-05-10-vibe-coding/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Here are the top messages from the talk:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What is vibe coding&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s where we ask the model to write &amp;amp; run code, don&amp;rsquo;t read the code, just inspect the &lt;strong&gt;behaviour&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s a &lt;strong&gt;coder&amp;rsquo;s tactic&lt;/strong&gt;, not a methodology. Use it when speed trumps certainty.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it&amp;rsquo;s catching on&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Non-coders can now ship apps&lt;/strong&gt; - no mental overhead of syntax or structure.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Coders think at a higher level&lt;/strong&gt; - stay in problem space, not bracket placement.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model capability keeps widening&lt;/strong&gt; - the &amp;ldquo;vibe-able&amp;rdquo; slice grows every release.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;How to work with it day-to-day&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Fail fast, hop models&lt;/strong&gt; - if Claude errors, paste into Gemini or OpenAI and move on.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Don&amp;rsquo;t fight sandbox limits&lt;/strong&gt; - browser LLM sandboxes block net calls; accept &amp;amp; upload files instead.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cross-validate outputs&lt;/strong&gt; - ask a second LLM to critique or replicate; cheaper than reading 400 lines of code.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Switch modes deliberately&lt;/strong&gt; - &lt;strong&gt;Vibe coding&lt;/strong&gt; when you don&amp;rsquo;t care about internals and time is scarce, &lt;strong&gt;AI-assisted coding&lt;/strong&gt; when you must own the code (read + tweak), &lt;strong&gt;Manual&lt;/strong&gt; only for the gnarly 5 % the model still can&amp;rsquo;t handle.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;What should we watch out for&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Security risk&lt;/strong&gt; - running unseen code can nuke your files; sandbox or use throw-away environments.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Internet-blocked runtimes&lt;/strong&gt; - prevents scraping/DoS misuse but forces data uploads.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Quality cliffs&lt;/strong&gt; - small edge-cases break; be ready to drop to manual fixes or wait for next model upgrade.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;What are the business implications&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Agencies still matter&lt;/strong&gt; - they absorb legal risk, project-manage, and can be bashed on price now that AI halves their grunt work.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prototype-to-prod blur&lt;/strong&gt; - the same vibe-coded PoC can often be hardened instead of rewritten.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;UI convergence&lt;/strong&gt; - chat + artifacts/canvas is becoming the default &amp;ldquo;front-end&amp;rdquo;; underlying apps become API + data.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;How does this impact education&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Curriculum can refresh term-by-term&lt;/strong&gt; - LLMs draft notes, slides, even whole modules.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Assessment shifts back to subjective&lt;/strong&gt; - LLM-graded essays/projects at scale.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Teach &amp;ldquo;learning how to learn&amp;rdquo;&lt;/strong&gt; - Pomodoro focus, spaced recall, chunking concepts, as in &lt;strong&gt;Learn Like a Pro&lt;/strong&gt; (Barbara Oakley).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Best tactic for staying current&lt;/strong&gt; - experiment &amp;gt; read; anything written is weeks out-of-date.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;What are the risks&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Overconfidence risk&lt;/strong&gt; - silent failures look like success until they hit prod.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Skill atrophy&lt;/strong&gt; - teams might lose the muscle to debug when vibe coding stalls.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Legal &amp;amp; compliance gaps&lt;/strong&gt; - unclear licence chains for AI-generated artefacts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Waiting game trap&lt;/strong&gt; - &amp;ldquo;just wait for the next model&amp;rdquo; can become a habit that freezes delivery.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href=&#34;https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7330549070744223745&#34;&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
  </channel>
</rss>
