<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>evaluation on S Anand</title>
    <link>https://www.s-anand.net/blog/tag/evaluation/</link>
    <description>Recent content in evaluation on S Anand</description>
    <generator>Hugo -- 0.164.0</generator>
    <language>en-us</language>
    <lastBuildDate>Tue, 24 Mar 2026 17:06:02 +0800</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/tag/evaluation/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Sonnet 4.6 vs MiniMax M2.7</title>
      <link>https://www.s-anand.net/blog/sonnet-4-6-vs-minimax-m2-7/</link>
      <pubDate>Tue, 24 Mar 2026 17:06:02 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/sonnet-4-6-vs-minimax-m2-7/</guid>
      <description>&lt;p&gt;Based on several (i.e. two) recommendations, I subscribed to &lt;a href=&#34;https://platform.minimax.io/&#34;&gt;MiniMax&lt;/a&gt;. At $10/month, you get 1,500 requests every 5 hours and 15,000 every week. That&amp;rsquo;s a LOT!&lt;/p&gt;
&lt;p&gt;Using the &lt;a href=&#34;https://sanand0.github.io/talks/2025-07-18-tug-true-but-irrelevant-rob-schrauwen/prompts.md&#34;&gt;same prompt&lt;/a&gt; I had &lt;a href=&#34;https://platform.minimax.io/docs/token-plan/claude-code&#34;&gt;Claude Code&lt;/a&gt; generate two data stories:&lt;/p&gt;
&lt;a href=&#34;https://sanand0.github.io/talks/2025-07-18-tug-true-but-irrelevant-rob-schrauwen/index.html&#34;&gt;
  &lt;figure&gt;
    &lt;img src=&#34;https://files.s-anand.net/images/2026-03-24-rob-data-story-claude.avif&#34; alt=&#34;In July 2025, in the heavy morning warmth of Thiruvananthapuram, the opening talk at the TeX User Group annual conference began with a crayon drawing. Not a polished diagram. Not a carefully curated slide. A hand-drawn family tree on crumpled paper, made by a six-year-old girl named Emily, who had a problem to solve. She had told her teacher she had seven grandmothers. The teacher declared it impossible. Emily went home and built a proof.&#34; /&gt;
    &lt;figcaption&gt;The first paragraph, by Claude Sonnet 4.6&lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/a&gt;
&lt;a href=&#34;https://sanand0.github.io/talks/2025-07-18-tug-true-but-irrelevant-rob-schrauwen/minimax.html&#34;&gt;
  &lt;figure&gt;
    &lt;img src=&#34;https://files.s-anand.net/images/2026-03-24-rob-data-story-minimax.avif&#34; alt=&#34;There is a drawing by a six-year-old girl in Amsterdam that contains more wisdom about data quality than three decades of enterprise content architecture. In it, seven figures stand in a family tree — each grandmother numbered with a circle, each grandfather conspicuously absent. The girl is Rob Schrauwen&#39;s granddaughter. She was asked to prove she had seven grandmothers; she drew them all, labeled them carefully, and in doing so invented continuous data quality, unique identifiers, and a knowledge graph — all before she learned to spell &amp;quot;Hetty.&amp;quot;&#34; /&gt;
    &lt;figcaption&gt;The first paragraph, by MiniMax M2.7&lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/a&gt;
&lt;p&gt;Here&amp;rsquo;s my comparison of the two. It&amp;rsquo;s partly based on &lt;a href=&#34;https://sanand0.github.io/talks/2025-07-18-tug-true-but-irrelevant-rob-schrauwen/comparison.md&#34;&gt;Claude Opus 4.6&amp;rsquo;s comparison&lt;/a&gt; but I felt the same way.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Dimension&lt;/th&gt;
					&lt;th&gt;Sonnet 4.6&lt;/th&gt;
					&lt;th&gt;MiniMax M2.7&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Narrative quality&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Immersive&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Content coverage&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Comprehensive&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Visual design&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;More varied, ambitious bands, no errors&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;CSS&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
					&lt;td&gt;Better use of CSS variables&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Tooltips&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Richer, comprehensive, &lt;code&gt;data-tip&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Modals/popups&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Richer, more types, more details&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Animated SVGs&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Richer, visually distinctive, sophisticated&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Slides&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Larger readable grid&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Code samples&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;XML vs JSON-LD side-by-side&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;External references&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Far more authoritative links&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Accessibility&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;ARIA, keyboard, alt text&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Generation quality&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Clean, no Chinese character artifacts&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;In other words, Sonnet 4.6 is a &lt;em&gt;clear&lt;/em&gt; winner on nearly every dimension.&lt;/p&gt;
&lt;p&gt;But the cost factor is &lt;em&gt;too&lt;/em&gt; big a difference to ignore. It feels like a 10x difference. So the question probably is: what can I do with a &lt;em&gt;reasonably&lt;/em&gt; good model that can generate 10X the quantity at the same price?&lt;/p&gt;
&lt;p&gt;(To be fair, &lt;a href=&#34;https://openrouter.ai/openai/gpt-5.4-mini&#34;&gt;GPT 5.4 Mini at 75c/MTok&lt;/a&gt; and &lt;a href=&#34;https://openrouter.ai/google/gemini-3-flash-preview&#34;&gt;Gemini 3 Flash at 50c/MTok&lt;/a&gt; are not far from &lt;a href=&#34;https://openrouter.ai/minimax/minimax-m2.7&#34;&gt;MiniMax M2.7 at 30c/MTok&lt;/a&gt; - but their &lt;a href=&#34;https://arena.ai/leaderboard/code&#34;&gt;code quality&lt;/a&gt; seems lower. I generated a &lt;a href=&#34;https://sanand0.github.io/talks/2025-07-18-tug-true-but-irrelevant-rob-schrauwen/gpt-5.4-mini-xhigh.html&#34;&gt;Codex - GPT 5.4 Mini version&lt;/a&gt; and while it has fewer errors it has even less visual style and narrative quality.)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Computer use&lt;/strong&gt; feels like a candidate. I used &lt;a href=&#34;https://github.com/simonw/rodney&#34;&gt;Rodney&lt;/a&gt; to research what drives my LinkedIn reach &amp;amp; engagement, and update my &lt;a href=&#34;https://github.com/sanand0/scripts/blob/f08ffd11e221c5a9ef58d5da814aaad9985bd422/agents/linkedin-cdp/SKILL.md&#34;&gt;SKILL.md&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I could try experimenting with sub-agents, doing bulk analysis (e.g. of code, transcripts, images), data discovery, etc. The crux of these is parallelization - something I have not explored much.&lt;/p&gt;
&lt;p&gt;It looks like twe&amp;rsquo;re entering an era where there are two kinds of use cases: high-quality for the best models, large-scale for the cheap models. The question is: how do I make the most of both?&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;a href=&#34;https://github.com/sanand0/talks/tree/52ad2aa775cd4e0f1e0ad8e6199ce7754a2663ac/2025-07-18-tug-true-but-irrelevant-rob-schrauwen&#34;&gt;Source Code&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;UPDATE&lt;/strong&gt;: Cheap models (or at least MiniMax M2.7) may be far less useful than I thought. I used MiniMax M2.7 with Claude Code for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;24 Mar 2026: Email analysis. I had it review my 15-year Gramener email archive for key events for a book. But it fetched too few results, so I switched to Codex (GPT 5.4 xhigh). &lt;!-- claude --resume 1c3edc84-5781-4684-bd0d-565fafc5b2b9 --&gt;&lt;/li&gt;
&lt;li&gt;25 Mar 2026: &lt;a href=&#34;https://play.picoctf.org/practice&#34;&gt;Capture The Flag&lt;/a&gt;. But it couldn&amp;rsquo;t solve problems, so I switched to Codex (GPT 5.4 xhigh). &lt;!-- claude --resume 525d2631-cf8b-4665-bfdc-e55b70cb7340 --&gt;&lt;/li&gt;
&lt;li&gt;25 Mar 2026: Songs download. I had it find popular Tamil songs and download them from YouTube. But the metadata was poor, so I switched to my own song collection. &lt;!-- claude --resume f4d3a74b-0bc8-4431-8084-f56be44a4a53 --&gt;&lt;/li&gt;
&lt;li&gt;26 Mar 2026: LEAN proofs. It started making too many basic mistakes (spelling errors in code!) I switched to Copilot (GPT 5.4 xhigh). &lt;!-- claude --resume a0836a15-6b82-45ef-8174-bcd10272a62e --&gt;&lt;/li&gt;
&lt;li&gt;29 Mar 2026: Calvin &amp;amp; Hobbes image analysis. It couldn&amp;rsquo;t even read the images and confidently saw &amp;ldquo;Hobbes stuck to a baseball bat with Mom &amp;amp; Dad&amp;rdquo; in a strip that only featured Calvin &amp;amp; Susie. &lt;!-- claude --resume 9dd15966-0c65-49bf-af3b-506f4bab5d39 --&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The main problems are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;It errs confidently&lt;/strong&gt;. It doesn&amp;rsquo;t do ROT13 well. It can&amp;rsquo;t see images. It mis-understands error messages. It assigned my earlier company&amp;rsquo;s incorporation date (NGIMAGE) as Gramener&amp;rsquo;s. It made Vijay Sethupathi a lyricist. When a process failed with just 12% coverage, it just continued. It just &lt;em&gt;reported what&amp;rsquo;s done, not what&amp;rsquo;s missing&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It&amp;rsquo;s a slow learner&lt;/strong&gt;. For &lt;a href=&#34;https://picoctf.org/&#34;&gt;picoCTF&lt;/a&gt;, it had the pieces but couldn&amp;rsquo;t assemble them. Claude Code resets the cwd, but it never switched to absolute paths. It mixed &lt;code&gt;uv run&lt;/code&gt; with &lt;code&gt;python3&lt;/code&gt;. It rewrites, resets or waits instead of diagnosing.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It&amp;rsquo;s best for simple, single-step tasks. Not where knowledge, accuracy, research matters. When using it, keep tasks small and verify correctness, completeness.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>IIM Bangalore PGP Interview Panel</title>
      <link>https://www.s-anand.net/blog/iim-bangalore-pgp-interview-panel/</link>
      <pubDate>Tue, 17 Mar 2026 06:42:32 +0530</pubDate>
      <guid>https://www.s-anand.net/blog/iim-bangalore-pgp-interview-panel/</guid>
      <description>&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-17-iim-bangalore-pgp-interview-panel.avif&#34;&gt; &lt;!-- https://gemini.google.com/u/2/app/2ff4c1dd6c10b1d7 --&gt;&lt;/p&gt;
&lt;p&gt;Yesterday, I was part of an IIM Bangalore interview panel at Hyderabad, along with Professor Subhabrata Das and Debajyoti. Panels typically comprise of two faculty and an alumni, and handle 8 interviews in the morning and eight in the evening, though in our case, we had 9 each.&lt;/p&gt;
&lt;p&gt;As we arrived, we were given a USB drive with the student&amp;rsquo;s resume, statement of purpose, and other documents that they had submitted, which included employment contracts, declarations, letters of recommendation, etc., depending on the student. Each interview was approximately 20 minutes. Luckily, Dr Das set a timer for 18, so we didn&amp;rsquo;t go too far beyond.&lt;/p&gt;
&lt;p&gt;There was practically no time to read the documents before each interview, so I used Claude with the following prompt to suggest questions.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-markdown&#34; data-lang=&#34;markdown&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;In the context of the below, suggest deep, probing questions that will reveal the suitability of the attached candidate for IIM PGP.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Feel free to search online for additional information about the candidate.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&amp;hellip;followed by&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The instructions email from IIMB&lt;/li&gt;
&lt;li&gt;The attachment in the email&lt;/li&gt;
&lt;li&gt;Claude&amp;rsquo;s earlier advice on how to evaluate candidates based on the above. This included suggestions to:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Probe for internal consistency&lt;/strong&gt;: &amp;ldquo;Your SoP says X but your job history suggests Y. Explain.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Probe for authenticity&lt;/strong&gt;: stuff that coaching can&amp;rsquo;t teach. (It gave candidate-specific questions.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Probe for social sensitivity&lt;/strong&gt;: &amp;ldquo;You&amp;rsquo;re from village X. Where does your aspiration help / hurt whihc segments of your village?&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The candidate&amp;rsquo;s documents (resume, statement of purpose, etc.)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;It helped that there were three panelists. That gave me time to screen the questions and understand the good ones - while others asked their questions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DISADVANTAGE&lt;/strong&gt;: You have to spend time reading the questions. Not all questions are great. Hallucinations exist (It said one candidate had a claw-back in their employment contract - they didn&amp;rsquo;t.)&lt;br&gt;
&lt;strong&gt;ADVANTAGE&lt;/strong&gt;: You pick up little-known stuff. One wrote a monthly salary instead of yearly and didn&amp;rsquo;t know it. The gaps in the resume, the dips in transcripts, etc. surfaced instantly.&lt;/p&gt;
&lt;p&gt;Initially, I just read out some questions.&lt;br&gt;
Then I tried using my own judgement.&lt;br&gt;
Then I went back to Claude&amp;rsquo;s questions.&lt;br&gt;
It took a while to learn how to use it well.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The interview questions idea came from Claude. It also suggested:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Transcribe real time and pop-up questions&lt;/strong&gt;. I wish I could and I&amp;rsquo;m sure it&amp;rsquo;ll happen soon.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Review scores &amp;amp; transcript with Claude&lt;/strong&gt;. I did almost exactly that!&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I told GitHub Copilot CLI (I had credits) running Claude 4.6 Sonnet (a sensible model):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;You are evaluating candidates for the PGP Program at IIM Bangalore.&lt;br&gt;
Read interview-guidelines.md fully to understand the process and criteria.&lt;/p&gt;
&lt;p&gt;Then, go through each candidate&amp;rsquo;s folder. These contain:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;notes.md: interview notes&lt;/li&gt;
&lt;li&gt;documents.md: documents submitted (converted from the PDF in that folder).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Primarily based on the notes (using documents.md for reference - no need to read PDFs), evaluate all candidates on the criteria mentioned in interview-guidelines.md.&lt;/p&gt;
&lt;p&gt;Plan like an expert interview evaluator first. In this context, first think about:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What patterns would an expert in this field check / recognize that beginners would miss?&lt;/li&gt;
&lt;li&gt;What questions would an expert ask that a beginner would not know to?&lt;/li&gt;
&lt;li&gt;What problems / failures would an expert anticipate that beginners may not be aware of?&lt;/li&gt;
&lt;li&gt;How would an expert analyze this? At each step, explain what they are looking for and why.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Document and plan your process in eval-copilot-claude/plan.md.&lt;br&gt;
Then, analyze students (use sub-agents as required) and document your analysis and evaluation in eval-copilot-claude/evaluation.md.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It reviewed all 18 candidates summarizing their background, work-experience, and interview scores along with strengths, concerns, and notable moments, with an evaluation summary.&lt;/p&gt;
&lt;p&gt;This serves as a good cross-check, I think. So I told it to:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Compare the ratings against my ratings (and notes) in README.md.&lt;br&gt;
Sort the candidates based on the difference between my ratings vs your (normalized / scaled from 1-10) ratings.&lt;br&gt;
For the biggest differences, analyze the reasons for the differences.&lt;br&gt;
Specifically, what might I have missed in the candidate that I should take a closer look at?&lt;br&gt;
Append this to eval-copilot-claude/evaluation.md.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I ran this on Copilot as well as Claude. We agreed on the best candidate. No question about that. Interesting!&lt;/p&gt;
&lt;p&gt;This revealed some interesting observations. I seem to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Ignore social insensitivity and artistic ability&lt;/strong&gt;. Most likely because I lack both. Note to self: I&amp;rsquo;m blind to, and hence devalue, what I&amp;rsquo;m not good at.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Value genuineness far more than analytics&lt;/strong&gt;. I rated one candidate with &lt;em&gt;terrible&lt;/em&gt; analysis skills as only a marginal reject. Claude rejected them outright. This is true for another candidate, too. But I rejected fakers outright.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;I calibrate for preparation&lt;/strong&gt;. I expect more from those with more experience, with more interview preparation, etc. Not a bad thing - just an observation.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;p&gt;In short:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;AI is great for preparation&lt;/strong&gt;. I should make it easier to feed it context.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI is great at post-mortems&lt;/strong&gt;. I need to budget time for this.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI can help &lt;em&gt;during&lt;/em&gt; discussions&lt;/strong&gt;. I need to figure out the technology for this.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Cracking online exams with coding agents</title>
      <link>https://www.s-anand.net/blog/cracking-online-exams-with-coding-agents/</link>
      <pubDate>Fri, 13 Mar 2026 15:37:19 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/cracking-online-exams-with-coding-agents/</guid>
      <description>&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-13-cracking-online-exams-with-coding-agents.avif&#34;&gt; &lt;!-- https://gemini.google.com/u/2/app/a6a6c341434ec848 --&gt;&lt;/p&gt;
&lt;p&gt;An effective way to solve online exams is to point a coding agent at it.&lt;/p&gt;
&lt;p&gt;I use that on my &lt;a href=&#34;https://tds.s-anand.net/&#34;&gt;Tools in Data Science&lt;/a&gt; course in two ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;As a test case of my code&lt;/strong&gt;. If my agent can solve it, good: I set the question correctly.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;As a test of student ability&lt;/strong&gt;. If it can&amp;rsquo;t, good: it&amp;rsquo;s a tough question (provided I didn&amp;rsquo;t make a mistake).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For &lt;a href=&#34;https://2026.pyconfhyd.org/&#34;&gt;PyConf, Hyderabad&lt;/a&gt;, my colleague built a &lt;a href=&#34;https://crack-the-prompt.straivedemo.com/&#34;&gt;Crack the Prompt&lt;/a&gt; challenge. Crack it and you get&amp;hellip; I don&amp;rsquo;t know&amp;hellip; goodies? A job interview? Leaderboard bragging rights?&lt;/p&gt;
&lt;p&gt;I told Codex:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Use the browser to visit &lt;a href=&#34;https://crack-the-prompt.straivedemo.com/&#34;&gt;https://crack-the-prompt.straivedemo.com/&lt;/a&gt; and solve it using the email ID &lt;a href=&#34;mailto:root.node@gmail.com&#34;&gt;root.node@gmail.com&lt;/a&gt; and GitHub handle sanand0&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;After 4 minutes, it told me:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The answers to all three prompt-engineering questions&lt;/li&gt;
&lt;li&gt;The code has a bug - so no one can submit &lt;em&gt;anyway&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;The prompts are hidden on the server-side (making it a bit harded to hack)&lt;/li&gt;
&lt;li&gt;But you can skip levels via the API - level-locking is front-end only&lt;/li&gt;
&lt;li&gt;&amp;hellip; and a whole bunch of interesting things.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;When I asked Claude to write about the process in Matt Levine&amp;rsquo;s style, &lt;a href=&#34;https://sanand0.github.io/datastories/crack-the-prompt/&#34;&gt;it included an interesting lesson&lt;/a&gt;.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The Victorians had the same problem. They designed elaborate entrance exams for the civil service because they wanted to identify people with the capacity for careful, systematic thinking. Then someone invented the civil service exam prep industry, and suddenly the exam was measuring preparation rather than capacity.&lt;/p&gt;
&lt;p&gt;The challenge was about the process &amp;ndash; about developing the instincts, the questioning strategies, the ability to read AI behavior like a poker tell. That&amp;rsquo;s the thing you can&amp;rsquo;t automate. Or rather, it&amp;rsquo;s the thing you can automate, which means it&amp;rsquo;s no longer a skill worth developing, which means we need to think about what skill we&amp;rsquo;re actually trying to cultivate when we design these challenges.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;EXACTLY&lt;/strong&gt;. What&amp;rsquo;s the skill we&amp;rsquo;re trying to cultivate? When will it be outdated?&lt;/p&gt;
&lt;p&gt;From now on, when testing, I&amp;rsquo;m going to write down &amp;ldquo;What skill is this &lt;strong&gt;really&lt;/strong&gt; testing?&amp;rdquo; That&amp;rsquo;s good enough a start.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>The Future of Work with AI</title>
      <link>https://www.s-anand.net/blog/the-future-of-work-with-ai/</link>
      <pubDate>Wed, 11 Mar 2026 13:56:58 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/the-future-of-work-with-ai/</guid>
      <description>&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-11-the-future-of-work-with-ai.avif&#34;&gt; &lt;!-- https://gemini.google.com/u/2/app/bad45f08e6f30526 --&gt;&lt;/p&gt;
&lt;!-- https://claude.ai/chat/6a2eb8d2-eba0-47a3-9447-a82a3336f62b --&gt;
&lt;p&gt;I often research how the world will change with AI by asking AI. Today&amp;rsquo;s session was informative. &lt;a href=&#34;https://claude.ai/share/18b521be-1820-494a-8df6-fde71b06fe6f&#34;&gt;I asked Claude, roughly&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Economics changes human behavior. As intelligence cost falls to zero, here are some changes in my behavior [I listed these].&lt;/p&gt;
&lt;p&gt;Others will have experienced behavioral changes too. Search online and synthesize behavioral changes.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It said this.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;🟡 People spend time on problem framing &amp;amp; evaluation. AI can execute the middle. (I&amp;rsquo;m OK at this. Need to do more framing + evaluation.)&lt;/li&gt;
&lt;li&gt;🟢 People don&amp;rsquo;t plan, they just build. (I&amp;rsquo;m &lt;a href=&#34;https://sanand0.github.io/&#34;&gt;prototyping&lt;/a&gt; a lot.)&lt;/li&gt;
&lt;li&gt;🟢 People build personal data &amp;amp; context. (I&amp;rsquo;m mining my &lt;a href=&#34;https://www.s-anand.net/blog/digital-exhaust/&#34;&gt;digital exhaust&lt;/a&gt;.)&lt;/li&gt;
&lt;li&gt;🔴 People queue work for agents, delegating into the future. (I&amp;rsquo;m not. I need to do &lt;strong&gt;far&lt;/strong&gt; more of this.)&lt;/li&gt;
&lt;li&gt;🟢 People shift from searching to asking for answers. (I do this a lot, e.g. this post.)&lt;/li&gt;
&lt;li&gt;🟡 People are AI-delegating junior jobs and developing senior level taste early. (Need to do more.)&lt;/li&gt;
&lt;li&gt;🟡 People treat unresolved emotions as prompts. (Need to do more.)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Rough legend: 🟢 = Stuff I know. 🟡 = I kind-of know. 🔴 = New learning.&lt;/p&gt;
&lt;p&gt;My next question was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Using economics, psychology, sociology, etc. &lt;strong&gt;predict new behavior changes&lt;/strong&gt; and implications.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It said: the bottleneck shift to trust, taste, and attention.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;🟢 &lt;strong&gt;Taste is a skill&lt;/strong&gt;. &lt;strong&gt;Few&lt;/strong&gt; people will deepen craft engagement (reading, cooking, music, &amp;hellip;) but most won&amp;rsquo;t, creating a gap.
&lt;ul&gt;
&lt;li&gt;Fuzzier areas (e.g. art, philosophy, unnamed ones) have more value than verifiable ones (e.g. code, science) since they&amp;rsquo;re harded to automate.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;🟢 &lt;strong&gt;Context is an asset&lt;/strong&gt;. People &amp;amp; companies will record decisions, reasoning, relationship histories, .. to create (marketplace-tradeable) assets.&lt;/li&gt;
&lt;li&gt;🟡 &lt;strong&gt;Trust is an asset&lt;/strong&gt;. People who stand behind AI output will create value in law, medicine, education, etc. Consulting will restructure around this. This will &lt;a href=&#34;https://arxiv.org/abs/2602.20946&#34;&gt;require investments&lt;/a&gt; &lt;a href=&#34;https://jack-clark.net/2026/03/02/import-ai-447-the-agi-economy-testing-ais-with-generated-games-and-agent-ecologies/&#34;&gt;#&lt;/a&gt; in:
&lt;ul&gt;
&lt;li&gt;Observability&lt;/li&gt;
&lt;li&gt;Human augmentation - gracefully degrading without harm when humans inevitably falter&lt;/li&gt;
&lt;li&gt;Synthetic practice&lt;/li&gt;
&lt;li&gt;Cryptographic provenance&lt;/li&gt;
&lt;li&gt;liability regimes that internalize tail risk&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;🟢 &lt;strong&gt;AI-less is exercise&lt;/strong&gt;. &lt;strong&gt;Few&lt;/strong&gt; people will deliberately train without AI, and will be valuable in niches (sports, emergencies, status-symbols). Exams / filters will partially screen for this.&lt;/li&gt;
&lt;li&gt;🟡 &lt;strong&gt;AI-less feels broken&lt;/strong&gt;. AI-speed will be &lt;em&gt;expected&lt;/em&gt; in services and interactions. &lt;strong&gt;Relationships are assets&lt;/strong&gt; that become valuable and premium.&lt;/li&gt;
&lt;li&gt;🔴 &lt;strong&gt;Orgs re-form&lt;/strong&gt;. Org sizes fall. But size has advantages (trust, brand, connects) so new orgs will form with teams evaluated on outcome (portfolio, ROI) and trust (network, ratings).&lt;/li&gt;
&lt;li&gt;🔴 &lt;strong&gt;Experience becomes luxury&lt;/strong&gt;. Non-reproducible experiences become expensive. Provable authenticity commands a premium.&lt;/li&gt;
&lt;li&gt;🔴 &lt;strong&gt;Two-speed world&lt;/strong&gt;. Some places (countries, companies, colleges, communities) become more AI friendly. Capital, talent and productivity will concentrate here.&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;What implication will these have on the nature of work?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Historically, apprenticeship (execution) preceded mastery (judgement). Now, execution is free. That messes things.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;🟡 &lt;strong&gt;How do we develop judgement?&lt;/strong&gt; Simulators, red-teaming, AI-free tests, &amp;hellip;?&lt;/li&gt;
&lt;li&gt;🔴 &lt;strong&gt;What will managers do?&lt;/strong&gt; Less coordination &amp;amp; oversight. More judgement (validate output not process), motivation, accountability. Like a film director, not supervisor.&lt;/li&gt;
&lt;li&gt;🟡 &lt;strong&gt;How will we hire/pay?&lt;/strong&gt; More outcome-based hire/pay. More freelancing, portfolio or reputation based hiring. Long-term retainers reserved for trust.&lt;/li&gt;
&lt;li&gt;🔴 &lt;strong&gt;How will we describe our work?&lt;/strong&gt; Less about tasks (which AI does) and more about who you are, where you fit, what you contribute.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So work will organize around:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Trust&lt;/strong&gt;: context, judgement, accountability&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Presence&lt;/strong&gt;: caring, building with hands, performing, connecting&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Direction&lt;/strong&gt;: framing, evaluating, curating&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&amp;hellip; and less around &lt;strong&gt;translation&lt;/strong&gt; (execution, coordination, oversight).&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Tools in Data Science - Jan 2026</title>
      <link>https://www.s-anand.net/blog/tools-in-data-science-jan-2026/</link>
      <pubDate>Thu, 29 Jan 2026 07:33:22 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/tools-in-data-science-jan-2026/</guid>
      <description>&lt;p&gt;My &lt;a href=&#34;https://tds.s-anand.net/&#34;&gt;Tools in Data Science course&lt;/a&gt; is available publicly, with a few changes from last year.&lt;/p&gt;
&lt;p&gt;First, I &lt;strong&gt;removed all the content&lt;/strong&gt;! Last year, Claude generated teaching material using my prompts. But what&amp;rsquo;s the point? I might as well give students the prompts directly. They can tweak it to their needs.&lt;/p&gt;
&lt;p&gt;This time, TDS shares the &lt;strong&gt;questions&lt;/strong&gt; needed to learn a topic. Any AI will give you good answers.&lt;/p&gt;
&lt;p&gt;Second, it focuses on &lt;strong&gt;what AI does NOT do well&lt;/strong&gt;. Coding syntax? Who cares. Basic analysis? ChatGPT can do that. In fact, each question now has an &amp;ldquo;Ask AI&amp;rdquo; button that dumps the question into your favorite AI tool. Just paste the answer and move on.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-01-29-tools-in-data-science-jan-2026.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;But&lt;/strong&gt;, these questions (hopefully) teach you &lt;em&gt;where AI fails&lt;/em&gt;. Putting tools together, prompting well, debugging and evaluating the output, etc. These matter more.&lt;/p&gt;
&lt;p&gt;Third, it&amp;rsquo;s &lt;strong&gt;easier to audit&lt;/strong&gt;. Anyone can take the course, even outside IITM. You can join the &lt;a href=&#34;https://groups.google.com/g/tds-iitm&#34;&gt;public Google Group&lt;/a&gt; for announcements. All &lt;a href=&#34;https://github.com/sanand0/tools-in-data-science-public/discussions&#34;&gt;questions are discussed publicly on GitHub&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;So, check the new version out, learn, and please share feedback!&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Google AI Tools List</title>
      <link>https://www.s-anand.net/blog/google-ai-tools-list/</link>
      <pubDate>Sun, 25 Jan 2026 10:35:32 +0530</pubDate>
      <guid>https://www.s-anand.net/blog/google-ai-tools-list/</guid>
      <description>&lt;p&gt;Google has released a huge number of AI tools. Not all are useful, but some are quite powerful. Here&amp;rsquo;s a list of the tools &lt;a href=&#34;https://chatgpt.com/share/6975a939-0398-8003-beea-2bc4c32f8ba8&#34;&gt;ChatGPT&lt;/a&gt; could find.&lt;/p&gt;
&lt;!-- https://chatgpt.com/c/6975a48b-edec-8320-b452-6731c9aed916 --&gt;
&lt;p&gt;🟢 = I find it good. 🟡 = Not too impressive. 🔴 = Avoid.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Assistants, research, and knowledge work
&lt;ul&gt;
&lt;li&gt;🟢 &lt;a href=&#34;https://gemini.google.com/&#34;&gt;Gemini&lt;/a&gt; is Google&amp;rsquo;s main AI assistant app. Use it as a &lt;em&gt;meeting-prep copilot&lt;/em&gt;: paste the agenda + last email thread, ask for &amp;ldquo;3 likely objections + crisp rebuttals + 5 questions that sound like I did my homework.&amp;rdquo;
&lt;ul&gt;
&lt;li&gt;🟢 &lt;a href=&#34;https://gemini.google/overview/deep-research/&#34;&gt;Gemini Deep Research&lt;/a&gt; is Gemini&amp;rsquo;s agentic research mode that browses many sources (optionally your Gmail/Drive/Chat) and produces multi-page reports. Use it to build a &lt;em&gt;client brief with citations&lt;/em&gt; (market, competitors, risks), then reuse it for outreach or a deck outline.&lt;/li&gt;
&lt;li&gt;🟢 &lt;a href=&#34;https://gemini.google/overview/canvas/&#34;&gt;Gemini Canvas&lt;/a&gt; turns ideas (and Deep Research reports) into shareable artifacts like web pages, quizzes, and simple apps. Use it to convert a research report into an &lt;em&gt;interactive explainer page&lt;/em&gt; your team can share internally.&lt;/li&gt;
&lt;li&gt;🟢 &lt;a href=&#34;https://gemini.google/overview/agent/&#34;&gt;Gemini Agent&lt;/a&gt; is an experimental &amp;ldquo;do multi-step tasks for me&amp;rdquo; feature that can use connected apps (Gmail/Calendar/Drive/Keep/Tasks, plus Maps/YouTube). Use it to &lt;em&gt;plan a week of customer check-ins&lt;/em&gt;: &amp;ldquo;find stalled deals, draft follow-ups, propose times, and create calendar holds-show me before sending.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;🟢 &lt;a href=&#34;https://notebooklm.google.com/&#34;&gt;NotebookLM&lt;/a&gt; is a source-grounded research notebook: it answers from your uploaded sources and can generate Audio Overviews. Use it to turn a messy folder of PDFs into a &lt;em&gt;decision memo&lt;/em&gt; + an &amp;ldquo;AI podcast&amp;rdquo; you can listen to while walking.&lt;/li&gt;
&lt;li&gt;🟡 &lt;a href=&#34;https://pinpoint.google.com/&#34;&gt;Pinpoint&lt;/a&gt; (Journalist Studio) helps explore huge collections of docs/audio/images with entity extraction and search. Use it for &lt;em&gt;internal investigations / audit trails&lt;/em&gt;: upload contracts + emails, then trace every mention of a vendor and its linked people/locations.&lt;/li&gt;
&lt;li&gt;🟢 &lt;a href=&#34;https://www.google.com/search?udm=50&#34;&gt;Google AI Mode&lt;/a&gt; exposes experimental Search experiences (including AI Mode where available). Use it for &lt;em&gt;rapid competitive scans&lt;/em&gt;: run the same query set weekly and track what changed in the AI-generated summaries vs links.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://blog.google/technology/google-labs/project-mariner/&#34;&gt;Project Mariner&lt;/a&gt; is a Google Labs &amp;ldquo;agentic&amp;rdquo; prototype aimed at taking actions on your behalf in a supervised way. Use it to &lt;em&gt;prototype a real workflow&lt;/em&gt; (e.g., &amp;ldquo;collect pricing from 20 vendor pages into a table&amp;rdquo;) before you invest in automating it properly.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Workspace and &amp;ldquo;AI inside Google apps&amp;rdquo;
&lt;ul&gt;
&lt;li&gt;🟢 &lt;a href=&#34;https://workspace.google.com/&#34;&gt;Google Workspace with Gemini&lt;/a&gt; brings Gemini into Gmail/Docs/Sheets/Drive, etc. Use it to &lt;em&gt;turn a weekly leadership email&lt;/em&gt; into: (1) action items per owner, (2) a draft reply, and (3) a one-slide summary for your staff meeting.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://workspace.google.com/products/vids/&#34;&gt;Google Vids&lt;/a&gt; is Workspace&amp;rsquo;s AI-assisted video creation tool. Use it to convert a project update doc into a &lt;em&gt;2-3 minute narrated update video&lt;/em&gt; for stakeholders who don&amp;rsquo;t read long emails.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://edu.google.com/intl/ALL_in/ai/gemini-for-education/&#34;&gt;Gemini for Education&lt;/a&gt; packages Gemini for teaching/learning contexts. Use it to generate &lt;em&gt;differentiated practice&lt;/em&gt;: same concept, three difficulty levels + a rubric + common misconceptions.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Build: developer + agent platforms
&lt;ul&gt;
&lt;li&gt;🟢 &lt;a href=&#34;https://aistudio.google.com/&#34;&gt;Google AI Studio&lt;/a&gt; is the fast path to prototyping with Gemini models and tools. Use it to build a &lt;em&gt;&amp;ldquo;contract red-flagger&amp;rdquo;&lt;/em&gt;: upload a contract, extract clauses into structured JSON, and generate a risk report you can paste into your workflow.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://firebase.studio/&#34;&gt;Firebase Studio&lt;/a&gt; is a browser-based &amp;ldquo;full-stack AI workspace&amp;rdquo; with agents, unifying Project IDX into Firebase. Use it to ship a &lt;em&gt;real internal tool&lt;/em&gt; (auth + UI + backend) without local setup, then deploy with Firebase/Cloud Run.&lt;/li&gt;
&lt;li&gt;🟢 &lt;a href=&#34;https://jules.google/&#34;&gt;Jules&lt;/a&gt; is an autonomous coding agent that connects to your GitHub repo and works through larger tasks on its own. E.g. give it “upgrade dependencies, fix the failing tests, and open a PR with a clear changelog,” then review it like a teammate’s PR instead of doing the grind yourself.
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://jules.google/docs/cli/reference/&#34;&gt;Jules Tools (CLI)&lt;/a&gt; is a command-line interface for running and monitoring Jules from your terminal or CI. E.g. pipe a TODO list into “one task per session,” auto-run nightly maintenance (lint/format/test fixes), and have it open PRs you can batch-review in the morning&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://developers.google.com/jules/api&#34;&gt;Jules API&lt;/a&gt; lets you programmatically trigger Jules from other systems. E.g. when a build fails, your pipeline can call the API with logs + stack trace, have Jules propose a fix + tests, and post a PR link back into Slack/Linear for human approval&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://idx.dev/&#34;&gt;Project IDX &amp;gt; Firebase Studio&lt;/a&gt; is the transition site if you used IDX. Use it to keep your existing workspaces but move to the newer Studio flows (agents + Gemini assistance).&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://genkit.dev/&#34;&gt;Genkit&lt;/a&gt; is an open-source framework for building AI-powered apps (workflows, tool use, structured output) across providers. Use it to productionize an &lt;em&gt;agentic workflow&lt;/em&gt; (RAG + tools + eval) with a local debugging UI before deployment.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://stax.withgoogle.com/&#34;&gt;Stax&lt;/a&gt; is Google’s evaluation platform for LLM apps (prompts, models, and end-to-end behaviors), built to replace “vibe testing” with repeatable scoring. E.g. codify your product’s rubric (tone, factuality, refusal correctness, latency), run it against every prompt/model change, and block releases when key metrics regress&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://deepmind.google/models/synthid/&#34;&gt;SynthID&lt;/a&gt; is DeepMind’s watermarking approach for identifying AI-generated/altered content. E.g. in an org that publishes lots of content, watermark what your tools generate and use detection as part of provenance checks before external release
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://ai.google.dev/responsible/docs/safeguards/synthid&#34;&gt;SynthID Text&lt;/a&gt; is the developer-facing tooling/docs for watermarking and detecting LLM-generated text. E.g. watermark outbound “AI-assisted” customer emails and automatically route them for review if they’re about regulated topics&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://ai.google.dev/responsible&#34;&gt;Responsible Generative AI Toolkit&lt;/a&gt; is Google’s “safeguards” hub: watermarking, safety classifiers, and guidance to reduce abuse and failure modes. E.g. wrap your app with layered defenses (input filtering + output moderation + policy tests) so one jailbreak prompt doesn’t become a security incident&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://cloud.google.com/products/agent-builder&#34;&gt;Vertex AI Agent Builder&lt;/a&gt; is Google Cloud&amp;rsquo;s platform to build, deploy, and govern enterprise agents grounded in enterprise data. Use it to build a &lt;em&gt;customer-support agent&lt;/em&gt; that can read policy docs, query BigQuery, and write safe responses with guardrails.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://codeassist.google/&#34;&gt;Gemini Code Assist&lt;/a&gt; is Gemini in your IDE (and beyond) with chat, completions, and agentic help. Use it for &lt;em&gt;large refactors&lt;/em&gt;: ask it to migrate a module, generate tests, and propose PR-ready diffs with explanations.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://pair.withgoogle.com/tools/&#34;&gt;PAIR Tools&lt;/a&gt; is Google’s hub of practical tools for understanding/debugging ML behavior (especially interpretability and fairness). E.g. before launch, run “slice analysis + counterfactual edits + feature sensitivity” to find where the model breaks on real user subgroups&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://pair-code.github.io/lit/&#34;&gt;LIT (Learning Interpretability Tool)&lt;/a&gt; is an interactive UI for probing models on text/image/tabular data. E.g. debug prompt brittleness by comparing outputs across controlled perturbations (tense, style, sensitive attributes) and visualizing salience/attribution to see what the model is actually using&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://pair-code.github.io/what-if-tool/&#34;&gt;What-If Tool&lt;/a&gt; is a minimal-coding tool to probe model predictions and fairness. E.g. manually edit a single example into multiple “what-if” counterfactuals and see which feature flips the decision, then turn that into a targeted data collection plan&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://pair-code.github.io/facets/&#34;&gt;Facets&lt;/a&gt; helps you explore and visualize datasets to catch skew, outliers, and leakage early. E.g. audit a training set for missingness and subgroup imbalance, then fix data before you waste time “tuning your way out” of a data problem&lt;/li&gt;
&lt;li&gt;🟡 &lt;a href=&#34;https://github.com/google-gemini/gemini-cli&#34;&gt;Gemini CLI&lt;/a&gt; brings Gemini into the terminal with file ops, shell commands, and search grounding. Use it as a &lt;em&gt;repo-native &amp;ldquo;ops copilot&amp;rdquo;&lt;/em&gt;: &amp;ldquo;scan logs, find the regression, propose the patch, run tests, and summarize.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;🟡 &lt;a href=&#34;https://deepmind.google/products/antigravity/&#34;&gt;Antigravity&lt;/a&gt; (DeepMind) is positioned as an agentic development environment. Use it when you want &lt;em&gt;multiple agents running tasks in parallel&lt;/em&gt; (debugging, refactoring, writing tests) while you supervise.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://docs.cloud.google.com/gemini/docs/overview&#34;&gt;Gemini for Google Cloud&lt;/a&gt; is Gemini embedded across many Google Cloud products. Use it for &lt;em&gt;cloud incident triage&lt;/em&gt;: summarize logs, hypothesize root cause, and generate the Terraform/IaC fix.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Create: media, design, marketing, and &amp;ldquo;labs&amp;rdquo; tools
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://labs.google/&#34;&gt;Google Labs&lt;/a&gt; is the hub for many experiments (Mixboard, Opal, CC, Learn Your Way, Doppl, etc.). Use it as your &amp;ldquo;what&amp;rsquo;s new&amp;rdquo; page-many tools show up here before they become mainstream.&lt;/li&gt;
&lt;li&gt;🟡 &lt;a href=&#34;https://opal.google/&#34;&gt;Opal&lt;/a&gt; builds, edits, and shares AI mini-apps from natural language (with a workflow editor). Use it to create a &lt;em&gt;repeatable analyst tool&lt;/em&gt; (e.g., &amp;ldquo;take a company name &amp;gt; pull recent news &amp;gt; summarize risks &amp;gt; draft outreach&amp;rdquo;).&lt;/li&gt;
&lt;li&gt;🟡 &lt;a href=&#34;https://mixboard.google.com/projects&#34;&gt;Mixboard&lt;/a&gt; is an AI concepting canvas/board for exploring and refining ideas. Use it to run a &lt;em&gt;structured ideation sprint&lt;/em&gt;: generate 20 variants, cluster them, then turn the top 3 into crisp one-pagers.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://labs.google.com/pomelli/about/&#34;&gt;Pomelli&lt;/a&gt; is a Labs marketing/brand tool that can infer brand identity and generate on-brand campaign assets. Use it to produce a &lt;em&gt;month of consistent social posts&lt;/em&gt; from your website + a few product photos.&lt;/li&gt;
&lt;li&gt;🟡 &lt;a href=&#34;https://stitch.withgoogle.com/&#34;&gt;Stitch&lt;/a&gt; turns prompts/sketches into UI designs and code. Use it to go from a rough wireframe to &lt;em&gt;React/Tailwind starter code&lt;/em&gt; you can hand to an engineer the same day.&lt;/li&gt;
&lt;li&gt;🟡 &lt;a href=&#34;https://labs.google/fx/tools/flow/&#34;&gt;Flow&lt;/a&gt; is a Labs tool aimed at AI video/story production workflows (built around Google&amp;rsquo;s gen-media stack). Use it to create a &lt;em&gt;pitch sizzle reel&lt;/em&gt; quickly: consistent characters + scenes + a simple timeline.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://labs.google/fx/tools/whisk/&#34;&gt;Whisk&lt;/a&gt; is a Labs image tool focused on controllable remixing (subject/scene/style style workflows). Use it for &lt;em&gt;fast, art-directable moodboards&lt;/em&gt; when text prompting is too loose.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://labs.google/fx/tools/image-fx/&#34;&gt;ImageFX&lt;/a&gt; is Google Labs&amp;rsquo; image-generation playground. Use it to iterate &lt;em&gt;brand-safe visual directions&lt;/em&gt; quickly (e.g., generate 30 &amp;ldquo;hero image&amp;rdquo; variants, pick 3, then refine).&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://labs.google/fx/tools/video-fx/&#34;&gt;VideoFX&lt;/a&gt; is the Labs surface for generative video (Veo-powered). Use it to prototype &lt;em&gt;short looping video backgrounds&lt;/em&gt; for product pages or events.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://labs.google/fx/tools/music-fx/&#34;&gt;MusicFX&lt;/a&gt; is the Labs music generation tool. Use it to generate &lt;em&gt;royalty-free stems&lt;/em&gt; (intro/outro/ambient) for podcasts or product videos.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://labs.google/doppl&#34;&gt;Doppl&lt;/a&gt; is a Labs try-on style experiment/app. Use it to sanity-check &lt;em&gt;creative wardrobe ideas&lt;/em&gt; before you buy, or to mock up &amp;ldquo;virtual merch&amp;rdquo; looks for a campaign.&lt;/li&gt;
&lt;li&gt;🟢 &lt;a href=&#34;https://gemini.google/overview/storybook/&#34;&gt;Gemini Storybook&lt;/a&gt; creates illustrated stories. Use it to generate &lt;em&gt;custom reading material&lt;/em&gt; for a specific learner&amp;rsquo;s interests (and adjust reading level/style).&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://textfx.withgoogle.com/&#34;&gt;TextFX&lt;/a&gt; is a Labs-style writing creativity tool (wordplay, transformations, constraints). Use it to generate &lt;em&gt;10 distinct &amp;ldquo;hooks&amp;rdquo;&lt;/em&gt; for the same idea before you write the real piece.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://labs.google/gentype&#34;&gt;GenType&lt;/a&gt; is a Labs experiment for AI-generated alphabets/type. Use it to create &lt;em&gt;a distinctive event identity&lt;/em&gt; (custom letterforms) without hiring a type designer for a one-off.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Science, security, and &amp;ldquo;serious AI&amp;rdquo;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://alphafoldserver.com/&#34;&gt;AlphaFold Server&lt;/a&gt; provides AlphaFold structure prediction as a web service. Use it to test &lt;em&gt;protein/ligand interaction hypotheses&lt;/em&gt; before spending lab time or compute on deeper simulations.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://cloud.google.com/security/products/threat-intelligence&#34;&gt;Google Threat Intelligence&lt;/a&gt; uses Gemini to help analyze threats and triage signals. Use it to turn a noisy alert stream into a &lt;em&gt;prioritized, explainable threat narrative&lt;/em&gt; your SOC can act on.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Models
&lt;ul&gt;
&lt;li&gt;🟡 &lt;a href=&#34;https://deepmind.google/models/gemma/&#34;&gt;Gemma&lt;/a&gt; is DeepMind’s family of lightweight open models built from the same tech lineage as Gemini. E.g. run a small, controlled model inside your VPC for narrow tasks (classification, extraction, safety filtering) when sending data to hosted LLMs is undesirable&lt;/li&gt;
&lt;li&gt;🟡 &lt;a href=&#34;https://cloud.google.com/model-garden&#34;&gt;Model Garden&lt;/a&gt; is Vertex AI’s catalog to discover, test, customize, and deploy models from Google and partners. E.g. shortlist 3 candidate models, run the same eval set, then deploy the winner behind one standardized platform with enterprise controls&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://cloud.google.com/generative-ai-studio&#34;&gt;Vertex AI Studio&lt;/a&gt; is the Google Cloud console surface for prototyping and testing genAI (prompts, model customization) in a governed environment. E.g. keep “prompt versions + test sets + pass/fail criteria” together so experiments become auditable artifacts, not scattered chats&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://ai.google.dev/edge/model-explorer&#34;&gt;Model Explorer&lt;/a&gt; helps you visually inspect model graphs so you can debug conversion/quantization and performance issues. E.g. compare two quantization strategies and pinpoint exactly which ops caused a latency spike or accuracy drop before you deploy&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://ai.google.dev/edge&#34;&gt;Google AI Edge&lt;/a&gt; is the umbrella for building on-device AI (mobile/web) with ready-to-use APIs across vision, audio, text, and genAI. E.g. ship an offline, privacy-preserving feature (document classification or on-device summarization) so latency and data exposure don’t depend on the network
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://ai.google.dev/edge/ai-edge-portal&#34;&gt;Google AI Edge Portal&lt;/a&gt; benchmarks LiteRT models across many real devices so you don’t guess performance from one phone. E.g. test the same model on a spread of target devices and pick the smallest model/config that consistently hits your FPS/latency target&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://playground.tensorflow.org/&#34;&gt;TensorFlow Playground&lt;/a&gt; is an interactive sandbox for understanding neural networks. E.g. use it to teach or debug intuitions—show how regularization, feature interactions, or class imbalance changes decision boundaries in minutes&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://teachablemachine.withgoogle.com/&#34;&gt;Teachable Machine&lt;/a&gt; lets anyone train simple image/sound/pose models in the browser and export them. E.g. prototype an accessibility feature (custom gesture or sound trigger) fast, then export the model to a small web demo your stakeholders can try&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Directories (&amp;ldquo;where to discover the rest&amp;rdquo;)
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://deepmind.google/&#34;&gt;Google DeepMind Products &amp;amp; Models&lt;/a&gt; (Gemini, Veo, Astra, Genie, etc.)-best &amp;ldquo;canonical list&amp;rdquo; of what exists.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://labs.google/experiments?category=develop&#34;&gt;Google Labs Experiments directory&lt;/a&gt;-browse by category (develop/create/learn) to catch smaller experiments you didn&amp;rsquo;t know to search for.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://experiments.withgoogle.com/&#34;&gt;Experiments with Google&lt;/a&gt; is a gallery of interactive demos (many AI) that’s great for prompt/data literacy and workshop “aha” moments. E.g. curate 5 experiments as a hands-on “AI intuition lab” for your team so they learn failure modes by playing, not by reading docs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title></title>
      <link>https://www.s-anand.net/blog/llms-are-smarter-than-us/</link>
      <pubDate>Tue, 01 Jul 2025 06:46:10 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/llms-are-smarter-than-us/</guid>
      <description>&lt;p&gt;LLMs are smarter than us in many areas. How do we control them?&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s not a new problem.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;VC partners&lt;/strong&gt; evaluate deep-tech startups.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Science editors&lt;/strong&gt; review Nobel laureates.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Managers&lt;/strong&gt; manage specialist teams.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Judges&lt;/strong&gt; evaluate expert testimony.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Coaches&lt;/strong&gt; train Olympic athletes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;… and they manage and evaluate &amp;ldquo;smarter&amp;rdquo; outputs in &lt;em&gt;many&lt;/em&gt; ways:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Verify&lt;/strong&gt;. Check against an &amp;ldquo;answer sheet&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Checklist&lt;/strong&gt;. Evaluate against pre-defined criteria.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sampling&lt;/strong&gt;. Randomly review a subset.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gating&lt;/strong&gt;. Accept low-risk work. Evaluate critical ones.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Benchmark&lt;/strong&gt;. Compare against others.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Red-team&lt;/strong&gt;. Probe to expose hidden flaws.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Double-blind review&lt;/strong&gt;. Mask identity to curb bias.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reproduce&lt;/strong&gt;. Re-running gives the same output?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Consensus&lt;/strong&gt;. Ask many. Wisdom of crowds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Outcome&lt;/strong&gt;. Did it work in the real world?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For example, you can apply them to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Vibe coding&lt;/strong&gt;: Non-programmers might glance at lint checks (&lt;em&gt;Checklist&lt;/em&gt;) and see if it works (&lt;em&gt;Outcome&lt;/em&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM image designs&lt;/strong&gt;: Developers might check if a few images look good (&lt;em&gt;Sampling&lt;/em&gt;) and check a few marketers (&lt;em&gt;Consensus&lt;/em&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM news articles&lt;/strong&gt;: An journalist might run a &lt;em&gt;Checklist&lt;/em&gt;, a &lt;em&gt;Double-blind review&lt;/em&gt; with experts, and &lt;em&gt;Verify&lt;/em&gt; critical facts (&lt;em&gt;Gating&lt;/em&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You &lt;em&gt;already&lt;/em&gt; know many of these. You learnt them in Auditing. Statistics. Law. System controls. Policy analysis. Quality engineering. Clinical epidemiology. Investigative journalism. Design critique.&lt;/p&gt;
&lt;p&gt;Worth brushing up on these skills. They&amp;rsquo;re &lt;em&gt;more&lt;/em&gt; important in the AI era.&lt;/p&gt;
&lt;p&gt;ChatGPT: &lt;a href=&#34;https://chatgpt.com/share/6863733f-4ebc-800c-ad3f-a2b472d9e9ca&#34;&gt;https://chatgpt.com/share/6863733f-4ebc-800c-ad3f-a2b472d9e9ca&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2025-07-01-llms-are-smarter-than-us-linkedin.jpg&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7345704252889059331&#34;&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>How To Control Smarter Intelligences</title>
      <link>https://www.s-anand.net/blog/how-to-control-smarter-intelligences/</link>
      <pubDate>Tue, 01 Jul 2025 06:37:40 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/how-to-control-smarter-intelligences/</guid>
      <description>&lt;p&gt;&lt;img alt=&#34;How To Control Smarter Intelligences&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/ChatGPT-Image-Jul-1-2025-11_27_23-AM.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;LLMs are smarter than us in many areas. How do we manage them?&lt;/p&gt;
&lt;p&gt;This is not a new problem.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;VC partners&lt;/strong&gt; evaluate deep-tech startups.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Science editors&lt;/strong&gt; review Nobel laureates.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Managers&lt;/strong&gt; manage specialist teams.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Judges&lt;/strong&gt; evaluate expert testimony.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Coaches&lt;/strong&gt; train Olympic athletes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;… and they manage and evaluate &amp;ldquo;smarter&amp;rdquo; outputs in &lt;strong&gt;many&lt;/strong&gt; ways:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Verify&lt;/strong&gt;. Check against an &amp;ldquo;answer sheet&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Checklist&lt;/strong&gt;. Evaluate against pre-defined criteria.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sampling&lt;/strong&gt;. Randomly review a subset.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gating&lt;/strong&gt;. Accept low-risk work. Evaluate critical ones.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Benchmark&lt;/strong&gt;. Compare against others.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Red-team&lt;/strong&gt;. Probe to expose hidden flaws.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Double-blind review&lt;/strong&gt;. Mask identity to curb bias.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reproduce&lt;/strong&gt;. Re-running gives the same output?&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Consensus&lt;/strong&gt;. Aggregate multiple responses. Wisdom of crowds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Outcome&lt;/strong&gt;. Did it work in the real world?&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Vibe coding&lt;/strong&gt;: Non-programmers might glance at lint checks (&lt;strong&gt;Checklist&lt;/strong&gt;) and see if it works (&lt;strong&gt;Outcome&lt;/strong&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM image designs&lt;/strong&gt;: Developers might check if a few images look good (&lt;strong&gt;Sampling&lt;/strong&gt;) and check a few marketers (&lt;strong&gt;Consensus&lt;/strong&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM news articles&lt;/strong&gt;: An journalist might run a &lt;strong&gt;Checklist&lt;/strong&gt;, a &lt;strong&gt;Double-blind review&lt;/strong&gt; with experts, and &lt;strong&gt;Verify&lt;/strong&gt; critical facts (&lt;strong&gt;Gating&lt;/strong&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;You &lt;strong&gt;already&lt;/strong&gt; know many of these. You learnt them in Auditing. Statistics. Law. System controls. Policy analysis. Quality engineering. Clinical epidemiology. Investigative journalism. Design critique.&lt;/p&gt;
&lt;p&gt;Worth brushing up skills. They&amp;rsquo;re &lt;strong&gt;more&lt;/strong&gt; important in the AI era.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>It&#39;s so easy to fool yourself</title>
      <link>https://www.s-anand.net/blog/its-so-easy-to-fool-yourself/</link>
      <pubDate>Sun, 30 Mar 2025 01:45:05 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/its-so-easy-to-fool-yourself/</guid>
      <description>&lt;p&gt;&lt;img alt=&#34;It&amp;rsquo;s so easy to fool yourself&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/ChatGPT-Image-Mar-30-2025-09_41_10-AM.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;I asked 4 LLMs to &amp;ldquo;Suggest catchy rhyming slogans to promote journaling.&amp;rdquo; Here are the suggestions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://chatgpt.com/share/67e89aa8-f5c4-800c-ae1c-f1ba4caaf9a9&#34;&gt;GPT 4.5&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Write it right, day or night!&lt;/li&gt;
&lt;li&gt;Reflect, inspect, journaling perfect!&lt;/li&gt;
&lt;li&gt;Express to impress, journaling’s best!&lt;/li&gt;
&lt;li&gt;Clear your mind, write and unwind!&lt;/li&gt;
&lt;li&gt;Pen your zen, again and again!&lt;/li&gt;
&lt;li&gt;Daily pages, inner sages!&lt;/li&gt;
&lt;li&gt;Take a pause, write your cause!&lt;/li&gt;
&lt;li&gt;Emotions flow, journals grow!&lt;/li&gt;
&lt;li&gt;Write, ignite, your inner light!&lt;/li&gt;
&lt;li&gt;Ink it, think it, don’t just blink it!&lt;/li&gt;
&lt;li&gt;Pen your quest, ease your stress!&lt;/li&gt;
&lt;li&gt;From confusion to conclusion—journal your solution!&lt;/li&gt;
&lt;li&gt;Keep a book, take a look, watch your growth hook!&lt;/li&gt;
&lt;li&gt;Thoughts on paper, worries taper!&lt;/li&gt;
&lt;li&gt;Write to heal, make it real!&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://claude.ai/share/eca94562-ac1c-4e28-b0de-2070e3d331b7&#34;&gt;Claude 3.7 Sonnet&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Write it down, turn that frown upside down!&lt;/li&gt;
&lt;li&gt;Pen your thoughts, connect the dots.&lt;/li&gt;
&lt;li&gt;Journal every day, keep the stress away.&lt;/li&gt;
&lt;li&gt;Ink your mind, treasures you&amp;rsquo;ll find.&lt;/li&gt;
&lt;li&gt;Write and reflect, your future self will thank you direct.&lt;/li&gt;
&lt;li&gt;Scribble and doodle, your mind will stay in the noodle.&lt;/li&gt;
&lt;li&gt;Journal with flair, show yourself you care.&lt;/li&gt;
&lt;li&gt;Jot it in ink, give yourself time to think.&lt;/li&gt;
&lt;li&gt;Fill the page, release the rage.&lt;/li&gt;
&lt;li&gt;Write your story, capture the glory.&lt;/li&gt;
&lt;li&gt;Daily writing, life enlightening.&lt;/li&gt;
&lt;li&gt;Express in a line, your day will shine.&lt;/li&gt;
&lt;li&gt;Thoughts on paper, life feels safer.&lt;/li&gt;
&lt;li&gt;Write with heart, that&amp;rsquo;s the smart start.&lt;/li&gt;
&lt;li&gt;Journaling&amp;rsquo;s neat, makes life complete.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%2215N0T8VgGv3cqN3t6An5RwnSP2UbhJU8J%22%5D,%22action%22:%22open%22,%22userId%22:%22108207362455575853987%22,%22resourceKeys%22:%7B%7D%7D&amp;amp;usp=sharing&#34;&gt;Gemini 2.5 Pro&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Thoughts unclear? Write them down here.&lt;/li&gt;
&lt;li&gt;Let the ink help you think.&lt;/li&gt;
&lt;li&gt;Untangle your mind, one line at a time.&lt;/li&gt;
&lt;li&gt;To make things clear, keep your journal near.&lt;/li&gt;
&lt;li&gt;When your head&amp;rsquo;s in a knot, give journaling a shot.&lt;/li&gt;
&lt;li&gt;Feeling stressed? Put your pen to the test.&lt;/li&gt;
&lt;li&gt;Write it down, let worries drown.&lt;/li&gt;
&lt;li&gt;Ease your mind, leave troubles behind.&lt;/li&gt;
&lt;li&gt;Let your feelings take flight, just sit down and write.&lt;/li&gt;
&lt;li&gt;Turn the page on inner rage.&lt;/li&gt;
&lt;li&gt;Write to know, watch yourself grow.&lt;/li&gt;
&lt;li&gt;Explore your soul, take control.&lt;/li&gt;
&lt;li&gt;Pen and page, turn a new age.&lt;/li&gt;
&lt;li&gt;Discover your way, write every day.&lt;/li&gt;
&lt;li&gt;Know your heart? Make journaling your art.&lt;/li&gt;
&lt;li&gt;Capture the day in your own way.&lt;/li&gt;
&lt;li&gt;Make moments last, write down the past.&lt;/li&gt;
&lt;li&gt;Don&amp;rsquo;t let memories stray, write them today.&lt;/li&gt;
&lt;li&gt;Reflect and connect, gain self-respect.&lt;/li&gt;
&lt;li&gt;Life moves fast? Make memories last.&lt;/li&gt;
&lt;li&gt;Get in the flow, write and grow.&lt;/li&gt;
&lt;li&gt;Pen in hand, understand.&lt;/li&gt;
&lt;li&gt;Daily lines, clear designs.&lt;/li&gt;
&lt;li&gt;Just take note, stay afloat.&lt;/li&gt;
&lt;li&gt;Write your heart, make a fresh start.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://x.com/i/grok/share/tmAxa8bTOl0kEcau3jqmLDEPT&#34;&gt;Grok 3&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;Write your day, let thoughts play!&lt;/li&gt;
&lt;li&gt;Pen in hand, your mind expands!&lt;/li&gt;
&lt;li&gt;Journal your way, seize the day!&lt;/li&gt;
&lt;li&gt;Feel the groove, let words move!&lt;/li&gt;
&lt;li&gt;Scribe your soul, make it whole!&lt;/li&gt;
&lt;li&gt;Ink flows free, just be me!&lt;/li&gt;
&lt;li&gt;Page by page, unlock your sage!&lt;/li&gt;
&lt;li&gt;Write it down, wear the crown!&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;At first glance, GPT 4.5 didn’t impress me. Claude 3.7 Sonnet did. I also didn’t like Gemini 2.5 Pro, but Grok was great.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Grok 3 &amp;gt; Claude 3.7 Sonnet &amp;gt; Gemini 2.5 Pro &amp;gt; GPT 4.5.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;But it’s hard to compare a dozen quotes at once. So I made a small &lt;a href=&#34;https://tools.s-anand.net/quotesarena/&#34;&gt;quotes arena app&lt;/a&gt; to help me pick my favorites. It shows me random pairs of quotes and asks which I like more.&lt;/p&gt;
&lt;p&gt;To my surprise, after answering 30+ &amp;ldquo;games&amp;rdquo; in the arena, I found that based on my preferences:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Claude 3.7 Sonnet &amp;gt; Gemini 2.5 Pro &amp;gt; GPT 4.5 &amp;gt; Grok 3.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That was weird. I thought I liked Grok&amp;rsquo;s results a lot. I continued till I answered 50+ games. Then I found that:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Grok 3 &amp;gt; GPT 4.5 &amp;gt; Gemini 2.5 Pro &amp;gt; Claude 3.7 Sonnet.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That&amp;rsquo;s the &lt;strong&gt;exact&lt;/strong&gt; opposite of the previous result.&lt;/p&gt;
&lt;p&gt;Honestly, I&amp;rsquo;m depressed. I&amp;rsquo;ve learnt 3 things:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;I can&amp;rsquo;t judge stuff at a glance.&lt;/li&gt;
&lt;li&gt;But I think I can (especially with code.)&lt;/li&gt;
&lt;li&gt;Even when evaluating carefully, my preferences are unstable.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Nothing&lt;/strong&gt; has shaken my confidence more in recent times. I &lt;strong&gt;cannot&lt;/strong&gt; trust my judgement. I need written evals. Badly.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>LLMs still do not locate bounding boxes well</title>
      <link>https://www.s-anand.net/blog/llms-still-do-not-locate-bounding-boxes-well/</link>
      <pubDate>Fri, 01 Nov 2024 16:20:03 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/llms-still-do-not-locate-bounding-boxes-well/</guid>
      <description>&lt;p&gt;I sent an image to over a dozen LLMs that support vision, asking them:&lt;/p&gt;
&lt;blockquote class=&#34;wp-block-quote&#34;&gt;
&lt;p&gt;Detect objects in this 1280x720 px image and return their color and bounding boxes in pixels. Respond as a JSON object: {[label]: [color, x1, y1, x2, y2], …}&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;None of the models did a good-enough job. It looks like we have some time to go before LLMs become good at bounding boxes.&lt;/p&gt;
&lt;p&gt;I&#39;ve given them a subjective rating on a 1-5 scale below. &lt;/p&gt;
&lt;figure class=&#34;wp-block-table&#34;&gt;&lt;table class=&#34;has-fixed-layout&#34;&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Model&lt;/th&gt;&lt;th&gt;Positions&lt;/th&gt;&lt;th&gt;Sizes&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;gemini-1.5-flash-001&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;td&gt;🟢🟢🟢🟢🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;gemini-1.5-flash-8b&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;gemini-1.5-flash-002&lt;/td&gt;&lt;td&gt;🟢🟢🔴🔴🔴&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;gemini-1.5-pro-002&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;td&gt;🟢🟢🟢🟢🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;gpt-4o-mini&lt;/td&gt;&lt;td&gt;🟢🔴🔴🔴🔴&lt;/td&gt;&lt;td&gt;🟢🟢🔴🔴🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;gpt-4o&lt;/td&gt;&lt;td&gt;🟢🟢🟢🟢🔴&lt;/td&gt;&lt;td&gt;🟢🟢🟢🟢🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;chatgpt-4o-latest&lt;/td&gt;&lt;td&gt;🟢🟢🟢🟢🔴&lt;/td&gt;&lt;td&gt;🟢🟢🟢🟢🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;claude-3-haiku-20240307&lt;/td&gt;&lt;td&gt;🟢🔴🔴🔴🔴&lt;/td&gt;&lt;td&gt;🟢🟢🔴🔴🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;claude-3=5-sonnet-20241022&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;llama-3.2-11b-vision-preview&lt;/td&gt;&lt;td&gt;🔴🔴🔴🔴🔴&lt;/td&gt;&lt;td&gt;🔴🔴🔴🔴🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;llama-3.2-90b-vision-preview&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;qwen-2-vl-72b-instruct&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;td&gt;🟢🟢🔴🔴🔴&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;pixtral-12b&lt;/td&gt;&lt;td&gt;🟢🟢🔴🔴🔴&lt;/td&gt;&lt;td&gt;🟢🟢🟢🔴🔴&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt;&lt;/figure&gt;
&lt;p&gt;I used an &lt;a href=&#34;https://tools.s-anand.net/llmboundingbox/&#34;&gt;app I built for this&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Here is the original image along with the individual results.&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/shapes-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/gemini-1.5-flash-8b-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/gemini-1.5-flash-001-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/gemini-1.5-flash-002-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/gemini-1.5-pro-002-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/gpt-4o-mini-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/gpt-4o-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/chatgpt-4o-latest-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/claude-3-haiku-20240307-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/claude-3-5-sonnet-20241022-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/llama-3.2-11b-vision-preview-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/llama-3.2-90b-vision-preview-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/pixtral-12b-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/qwen-2-vl-72b-instruct-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;h4 class=&#34;wp-block-heading&#34;&gt;Update&lt;/h4&gt;
&lt;p&gt;Adding gridlines with labeled axes helps the LLMs. (Thanks &lt;a href=&#34;https://www.linkedin.com/feed/update/urn:li:ugcPost:7258152478255243264?commentUrn=urn%3Ali%3Acomment%3A%28ugcPost%3A7258152478255243264%2C7258184350876262402%29&amp;amp;dashCommentUrn=urn%3Ali%3Afsd_comment%3A%287258184350876262402%2Curn%3Ali%3AugcPost%3A7258152478255243264%29&#34;&gt;@Bijan Mishra&lt;/a&gt;.) Here are a few examples:&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/chatgpt-4o-latest-1-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/gemini-1.5-flash-002-1-1024x576.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/claude-3-5-sonnet-20241022-1-1024x576.webp&#34;&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>How does Gemini process videos?</title>
      <link>https://www.s-anand.net/blog/how-does-gemini-process-videos/</link>
      <pubDate>Thu, 24 Oct 2024 08:47:21 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/how-does-gemini-process-videos/</guid>
      <description>&lt;p&gt;The Gemini documentation is clear:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The File API service extracts image frames from videos at 1 frame per second (FPS) and audio at 1Kbps, single channel, adding timestamps every second. These rates are subject to change in the future for improvements in inference.&lt;/p&gt;
&lt;p&gt;Note: The details of fast action sequences may be lost at the 1 FPS frame sampling rate. Consider slowing down high-speed clips for improved inference quality.&lt;/p&gt;
&lt;p&gt;Individual frames are 258 tokens, and audio is 32 tokens per second. With metadata, each second of video becomes ~300 tokens, which means a 1M context window can fit slightly less than an hour of video.&lt;/p&gt;
&lt;p&gt;To ask questions about time-stamped locations, use the format MM:SS, where the first two digits represent minutes and the last two digits represent seconds.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;But on this &lt;a href=&#34;https://sub.thursdai.news/p/thursdai-oct-17-robots-rockets-and&#34;&gt;ThursdAI episode: Oct 17 - Robots, Rockets, and Multi Modal Mania&amp;hellip;&lt;/a&gt;, at 1:00:50, &lt;a href=&#34;https://x.com/hrishioa/&#34;&gt;Hrishi&lt;/a&gt; says&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I don&amp;rsquo;t think it&amp;rsquo;s a series of images anymore because when I talk to the model and try to get some concept of what it&amp;rsquo;s perceiving, it&amp;rsquo;s no longer a series of images.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;If that&amp;rsquo;s the case, it&amp;rsquo;s a &lt;strong&gt;huge&lt;/strong&gt; change. So I tested it with this video.&lt;/p&gt;
&lt;div class=&#34;video-embed&#34;&gt;&lt;iframe src=&#34;https://www.youtube.com/embed/Dv8KON7WQYA&#34; title=&#34;YouTube video&#34; loading=&#34;lazy&#34; allow=&#34;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture&#34; allowfullscreen&gt;&lt;/iframe&gt;&lt;/div&gt;
&lt;p&gt;This video has 20 numbers refreshing at 4 frames per second.&lt;/p&gt;
&lt;p&gt;When I upload it to &lt;a href=&#34;https://aistudio.google.com/&#34;&gt;AI Studio&lt;/a&gt;, it takes 1,316 tokens. This is close enough to 258 tokens per image (no audio). So I&amp;rsquo;m partly convinced that Gemini still processing videos at 1 frame per second.&lt;/p&gt;
&lt;p&gt;Then, I asked it to &lt;code&gt;Extract all numbers in the video&lt;/code&gt; using Gemini 1.5 Flash 002 as well as Gemini 1.5 Flash 8b. In both cases, the results were: 2018, 85, 47, 37, 38.&lt;/p&gt;
&lt;p&gt;These are frames 2, 6, 10, 14, 18 (out of 20). So, &lt;strong&gt;clearly&lt;/strong&gt; Gemini is still sampling at about 1 frame per second, starting somewhere between 0.25 or 0.5 seconds.&lt;/p&gt;
</description>
    </item>
  </channel>
</rss>
