<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>llm-comparison on S Anand</title>
    <link>https://www.s-anand.net/blog/tag/llm-comparison/</link>
    <description>Recent content in llm-comparison on S Anand</description>
    <generator>Hugo -- 0.164.0</generator>
    <language>en-us</language>
    <lastBuildDate>Tue, 24 Mar 2026 17:06:02 +0800</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/tag/llm-comparison/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Sonnet 4.6 vs MiniMax M2.7</title>
      <link>https://www.s-anand.net/blog/sonnet-4-6-vs-minimax-m2-7/</link>
      <pubDate>Tue, 24 Mar 2026 17:06:02 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/sonnet-4-6-vs-minimax-m2-7/</guid>
      <description>&lt;p&gt;Based on several (i.e. two) recommendations, I subscribed to &lt;a href=&#34;https://platform.minimax.io/&#34;&gt;MiniMax&lt;/a&gt;. At $10/month, you get 1,500 requests every 5 hours and 15,000 every week. That&amp;rsquo;s a LOT!&lt;/p&gt;
&lt;p&gt;Using the &lt;a href=&#34;https://sanand0.github.io/talks/2025-07-18-tug-true-but-irrelevant-rob-schrauwen/prompts.md&#34;&gt;same prompt&lt;/a&gt; I had &lt;a href=&#34;https://platform.minimax.io/docs/token-plan/claude-code&#34;&gt;Claude Code&lt;/a&gt; generate two data stories:&lt;/p&gt;
&lt;a href=&#34;https://sanand0.github.io/talks/2025-07-18-tug-true-but-irrelevant-rob-schrauwen/index.html&#34;&gt;
  &lt;figure&gt;
    &lt;img src=&#34;https://files.s-anand.net/images/2026-03-24-rob-data-story-claude.avif&#34; alt=&#34;In July 2025, in the heavy morning warmth of Thiruvananthapuram, the opening talk at the TeX User Group annual conference began with a crayon drawing. Not a polished diagram. Not a carefully curated slide. A hand-drawn family tree on crumpled paper, made by a six-year-old girl named Emily, who had a problem to solve. She had told her teacher she had seven grandmothers. The teacher declared it impossible. Emily went home and built a proof.&#34; /&gt;
    &lt;figcaption&gt;The first paragraph, by Claude Sonnet 4.6&lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/a&gt;
&lt;a href=&#34;https://sanand0.github.io/talks/2025-07-18-tug-true-but-irrelevant-rob-schrauwen/minimax.html&#34;&gt;
  &lt;figure&gt;
    &lt;img src=&#34;https://files.s-anand.net/images/2026-03-24-rob-data-story-minimax.avif&#34; alt=&#34;There is a drawing by a six-year-old girl in Amsterdam that contains more wisdom about data quality than three decades of enterprise content architecture. In it, seven figures stand in a family tree — each grandmother numbered with a circle, each grandfather conspicuously absent. The girl is Rob Schrauwen&#39;s granddaughter. She was asked to prove she had seven grandmothers; she drew them all, labeled them carefully, and in doing so invented continuous data quality, unique identifiers, and a knowledge graph — all before she learned to spell &amp;quot;Hetty.&amp;quot;&#34; /&gt;
    &lt;figcaption&gt;The first paragraph, by MiniMax M2.7&lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/a&gt;
&lt;p&gt;Here&amp;rsquo;s my comparison of the two. It&amp;rsquo;s partly based on &lt;a href=&#34;https://sanand0.github.io/talks/2025-07-18-tug-true-but-irrelevant-rob-schrauwen/comparison.md&#34;&gt;Claude Opus 4.6&amp;rsquo;s comparison&lt;/a&gt; but I felt the same way.&lt;/p&gt;
&lt;table&gt;
	&lt;thead&gt;
			&lt;tr&gt;
					&lt;th&gt;Dimension&lt;/th&gt;
					&lt;th&gt;Sonnet 4.6&lt;/th&gt;
					&lt;th&gt;MiniMax M2.7&lt;/th&gt;
			&lt;/tr&gt;
	&lt;/thead&gt;
	&lt;tbody&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Narrative quality&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Immersive&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Content coverage&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Comprehensive&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Visual design&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;More varied, ambitious bands, no errors&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;CSS&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
					&lt;td&gt;Better use of CSS variables&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Tooltips&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Richer, comprehensive, &lt;code&gt;data-tip&lt;/code&gt;&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Modals/popups&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Richer, more types, more details&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Animated SVGs&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Richer, visually distinctive, sophisticated&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Slides&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Larger readable grid&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Code samples&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;XML vs JSON-LD side-by-side&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;External references&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Far more authoritative links&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Accessibility&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;ARIA, keyboard, alt text&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
			&lt;tr&gt;
					&lt;td&gt;&lt;strong&gt;Generation quality&lt;/strong&gt;&lt;/td&gt;
					&lt;td&gt;Clean, no Chinese character artifacts&lt;/td&gt;
					&lt;td&gt;&lt;/td&gt;
			&lt;/tr&gt;
	&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;In other words, Sonnet 4.6 is a &lt;em&gt;clear&lt;/em&gt; winner on nearly every dimension.&lt;/p&gt;
&lt;p&gt;But the cost factor is &lt;em&gt;too&lt;/em&gt; big a difference to ignore. It feels like a 10x difference. So the question probably is: what can I do with a &lt;em&gt;reasonably&lt;/em&gt; good model that can generate 10X the quantity at the same price?&lt;/p&gt;
&lt;p&gt;(To be fair, &lt;a href=&#34;https://openrouter.ai/openai/gpt-5.4-mini&#34;&gt;GPT 5.4 Mini at 75c/MTok&lt;/a&gt; and &lt;a href=&#34;https://openrouter.ai/google/gemini-3-flash-preview&#34;&gt;Gemini 3 Flash at 50c/MTok&lt;/a&gt; are not far from &lt;a href=&#34;https://openrouter.ai/minimax/minimax-m2.7&#34;&gt;MiniMax M2.7 at 30c/MTok&lt;/a&gt; - but their &lt;a href=&#34;https://arena.ai/leaderboard/code&#34;&gt;code quality&lt;/a&gt; seems lower. I generated a &lt;a href=&#34;https://sanand0.github.io/talks/2025-07-18-tug-true-but-irrelevant-rob-schrauwen/gpt-5.4-mini-xhigh.html&#34;&gt;Codex - GPT 5.4 Mini version&lt;/a&gt; and while it has fewer errors it has even less visual style and narrative quality.)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Computer use&lt;/strong&gt; feels like a candidate. I used &lt;a href=&#34;https://github.com/simonw/rodney&#34;&gt;Rodney&lt;/a&gt; to research what drives my LinkedIn reach &amp;amp; engagement, and update my &lt;a href=&#34;https://github.com/sanand0/scripts/blob/f08ffd11e221c5a9ef58d5da814aaad9985bd422/agents/linkedin-cdp/SKILL.md&#34;&gt;SKILL.md&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I could try experimenting with sub-agents, doing bulk analysis (e.g. of code, transcripts, images), data discovery, etc. The crux of these is parallelization - something I have not explored much.&lt;/p&gt;
&lt;p&gt;It looks like twe&amp;rsquo;re entering an era where there are two kinds of use cases: high-quality for the best models, large-scale for the cheap models. The question is: how do I make the most of both?&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;a href=&#34;https://github.com/sanand0/talks/tree/52ad2aa775cd4e0f1e0ad8e6199ce7754a2663ac/2025-07-18-tug-true-but-irrelevant-rob-schrauwen&#34;&gt;Source Code&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;UPDATE&lt;/strong&gt;: Cheap models (or at least MiniMax M2.7) may be far less useful than I thought. I used MiniMax M2.7 with Claude Code for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;24 Mar 2026: Email analysis. I had it review my 15-year Gramener email archive for key events for a book. But it fetched too few results, so I switched to Codex (GPT 5.4 xhigh). &lt;!-- claude --resume 1c3edc84-5781-4684-bd0d-565fafc5b2b9 --&gt;&lt;/li&gt;
&lt;li&gt;25 Mar 2026: &lt;a href=&#34;https://play.picoctf.org/practice&#34;&gt;Capture The Flag&lt;/a&gt;. But it couldn&amp;rsquo;t solve problems, so I switched to Codex (GPT 5.4 xhigh). &lt;!-- claude --resume 525d2631-cf8b-4665-bfdc-e55b70cb7340 --&gt;&lt;/li&gt;
&lt;li&gt;25 Mar 2026: Songs download. I had it find popular Tamil songs and download them from YouTube. But the metadata was poor, so I switched to my own song collection. &lt;!-- claude --resume f4d3a74b-0bc8-4431-8084-f56be44a4a53 --&gt;&lt;/li&gt;
&lt;li&gt;26 Mar 2026: LEAN proofs. It started making too many basic mistakes (spelling errors in code!) I switched to Copilot (GPT 5.4 xhigh). &lt;!-- claude --resume a0836a15-6b82-45ef-8174-bcd10272a62e --&gt;&lt;/li&gt;
&lt;li&gt;29 Mar 2026: Calvin &amp;amp; Hobbes image analysis. It couldn&amp;rsquo;t even read the images and confidently saw &amp;ldquo;Hobbes stuck to a baseball bat with Mom &amp;amp; Dad&amp;rdquo; in a strip that only featured Calvin &amp;amp; Susie. &lt;!-- claude --resume 9dd15966-0c65-49bf-af3b-506f4bab5d39 --&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The main problems are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;It errs confidently&lt;/strong&gt;. It doesn&amp;rsquo;t do ROT13 well. It can&amp;rsquo;t see images. It mis-understands error messages. It assigned my earlier company&amp;rsquo;s incorporation date (NGIMAGE) as Gramener&amp;rsquo;s. It made Vijay Sethupathi a lyricist. When a process failed with just 12% coverage, it just continued. It just &lt;em&gt;reported what&amp;rsquo;s done, not what&amp;rsquo;s missing&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It&amp;rsquo;s a slow learner&lt;/strong&gt;. For &lt;a href=&#34;https://picoctf.org/&#34;&gt;picoCTF&lt;/a&gt;, it had the pieces but couldn&amp;rsquo;t assemble them. Claude Code resets the cwd, but it never switched to absolute paths. It mixed &lt;code&gt;uv run&lt;/code&gt; with &lt;code&gt;python3&lt;/code&gt;. It rewrites, resets or waits instead of diagnosing.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It&amp;rsquo;s best for simple, single-step tasks. Not where knowledge, accuracy, research matters. When using it, keep tasks small and verify correctness, completeness.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Which LLMs get you better grades?</title>
      <link>https://www.s-anand.net/blog/which-llms-get-you-better-grades/</link>
      <pubDate>Fri, 06 Mar 2026 19:26:47 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/which-llms-get-you-better-grades/</guid>
      <description>&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-06-which-llm-get-you-better-grades.avif&#34;&gt; &lt;!-- https://gemini.google.com/app/72f962e80615e800 --&gt;&lt;/p&gt;
&lt;p&gt;In my &lt;a href=&#34;https://exam.sanand.workers.dev/tds-2026-01-ee&#34;&gt;graded assignments&lt;/a&gt; students can pick an AI and &amp;ldquo;Ask AI&amp;rdquo; any question at the click of a button. It defaults to Google AI Mode, but other models are available. I know who uses which model and their scores in each assignment.&lt;/p&gt;
&lt;p&gt;I asked Codex to test the hypothesis whether using a specific model helps students perform better.&lt;/p&gt;
&lt;p&gt;The short answer? Yes. Model choice matters a lot. Across 333 students, here&amp;rsquo;s how much more/less students score compared with ChatGPT:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Perplexity:&lt;/strong&gt; increases scores by ~11% points (99.8% sure it has an impact).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude:&lt;/strong&gt; +9% (99.5% sure)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Clipboard:&lt;/strong&gt;, i.e. just copying to the clipboard and pasting in their choice of AI: +9% points (97.5% sure)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Google:&lt;/strong&gt; -1% (Not sure. Sorry, Sundar.)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you are still using default ChatGPT, you are leaving nearly a full letter grade on the table compared to Claude or Perplexity users.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;I also looked at exam timing. Do specific models perform better under last-minute pressure?&lt;/p&gt;
&lt;p&gt;Statistically, no. The model-by-timing interaction isn&amp;rsquo;t significant (92% sure). A good model won&amp;rsquo;t save a rushed submission any better than a bad one.&lt;/p&gt;
&lt;p&gt;But procrastination does hurt, regardless of which AI they frantically prompt. In the &lt;a href=&#34;https://exam.sanand.workers.dev/tds-2026-01-ga1&#34;&gt;first graded assignment (GA1)&lt;/a&gt;, there was a negative correlation (-37%, 100% sure) between submission time and final score.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;BTW, I correlated with the &lt;em&gt;most-used&lt;/em&gt; model. I can&amp;rsquo;t guarantee they used that specific model on every single attempt.&lt;/p&gt;
&lt;p&gt;Of course, correlation isn&amp;rsquo;t causation. Maybe Claude writes better code. Or perhaps the kind of student who consciously switches from ChatGPT to Perplexity or Claude is just a better, more engaged student.&lt;/p&gt;
&lt;p&gt;But a 9-to-10 point bump is a pretty good reason to experiment.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>It&#39;s so easy to fool yourself</title>
      <link>https://www.s-anand.net/blog/its-so-easy-to-fool-yourself/</link>
      <pubDate>Sun, 30 Mar 2025 01:45:05 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/its-so-easy-to-fool-yourself/</guid>
      <description>&lt;p&gt;&lt;img alt=&#34;It&amp;rsquo;s so easy to fool yourself&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/ChatGPT-Image-Mar-30-2025-09_41_10-AM.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;I asked 4 LLMs to &amp;ldquo;Suggest catchy rhyming slogans to promote journaling.&amp;rdquo; Here are the suggestions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://chatgpt.com/share/67e89aa8-f5c4-800c-ae1c-f1ba4caaf9a9&#34;&gt;GPT 4.5&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Write it right, day or night!&lt;/li&gt;
&lt;li&gt;Reflect, inspect, journaling perfect!&lt;/li&gt;
&lt;li&gt;Express to impress, journaling’s best!&lt;/li&gt;
&lt;li&gt;Clear your mind, write and unwind!&lt;/li&gt;
&lt;li&gt;Pen your zen, again and again!&lt;/li&gt;
&lt;li&gt;Daily pages, inner sages!&lt;/li&gt;
&lt;li&gt;Take a pause, write your cause!&lt;/li&gt;
&lt;li&gt;Emotions flow, journals grow!&lt;/li&gt;
&lt;li&gt;Write, ignite, your inner light!&lt;/li&gt;
&lt;li&gt;Ink it, think it, don’t just blink it!&lt;/li&gt;
&lt;li&gt;Pen your quest, ease your stress!&lt;/li&gt;
&lt;li&gt;From confusion to conclusion—journal your solution!&lt;/li&gt;
&lt;li&gt;Keep a book, take a look, watch your growth hook!&lt;/li&gt;
&lt;li&gt;Thoughts on paper, worries taper!&lt;/li&gt;
&lt;li&gt;Write to heal, make it real!&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://claude.ai/share/eca94562-ac1c-4e28-b0de-2070e3d331b7&#34;&gt;Claude 3.7 Sonnet&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Write it down, turn that frown upside down!&lt;/li&gt;
&lt;li&gt;Pen your thoughts, connect the dots.&lt;/li&gt;
&lt;li&gt;Journal every day, keep the stress away.&lt;/li&gt;
&lt;li&gt;Ink your mind, treasures you&amp;rsquo;ll find.&lt;/li&gt;
&lt;li&gt;Write and reflect, your future self will thank you direct.&lt;/li&gt;
&lt;li&gt;Scribble and doodle, your mind will stay in the noodle.&lt;/li&gt;
&lt;li&gt;Journal with flair, show yourself you care.&lt;/li&gt;
&lt;li&gt;Jot it in ink, give yourself time to think.&lt;/li&gt;
&lt;li&gt;Fill the page, release the rage.&lt;/li&gt;
&lt;li&gt;Write your story, capture the glory.&lt;/li&gt;
&lt;li&gt;Daily writing, life enlightening.&lt;/li&gt;
&lt;li&gt;Express in a line, your day will shine.&lt;/li&gt;
&lt;li&gt;Thoughts on paper, life feels safer.&lt;/li&gt;
&lt;li&gt;Write with heart, that&amp;rsquo;s the smart start.&lt;/li&gt;
&lt;li&gt;Journaling&amp;rsquo;s neat, makes life complete.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%2215N0T8VgGv3cqN3t6An5RwnSP2UbhJU8J%22%5D,%22action%22:%22open%22,%22userId%22:%22108207362455575853987%22,%22resourceKeys%22:%7B%7D%7D&amp;amp;usp=sharing&#34;&gt;Gemini 2.5 Pro&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Thoughts unclear? Write them down here.&lt;/li&gt;
&lt;li&gt;Let the ink help you think.&lt;/li&gt;
&lt;li&gt;Untangle your mind, one line at a time.&lt;/li&gt;
&lt;li&gt;To make things clear, keep your journal near.&lt;/li&gt;
&lt;li&gt;When your head&amp;rsquo;s in a knot, give journaling a shot.&lt;/li&gt;
&lt;li&gt;Feeling stressed? Put your pen to the test.&lt;/li&gt;
&lt;li&gt;Write it down, let worries drown.&lt;/li&gt;
&lt;li&gt;Ease your mind, leave troubles behind.&lt;/li&gt;
&lt;li&gt;Let your feelings take flight, just sit down and write.&lt;/li&gt;
&lt;li&gt;Turn the page on inner rage.&lt;/li&gt;
&lt;li&gt;Write to know, watch yourself grow.&lt;/li&gt;
&lt;li&gt;Explore your soul, take control.&lt;/li&gt;
&lt;li&gt;Pen and page, turn a new age.&lt;/li&gt;
&lt;li&gt;Discover your way, write every day.&lt;/li&gt;
&lt;li&gt;Know your heart? Make journaling your art.&lt;/li&gt;
&lt;li&gt;Capture the day in your own way.&lt;/li&gt;
&lt;li&gt;Make moments last, write down the past.&lt;/li&gt;
&lt;li&gt;Don&amp;rsquo;t let memories stray, write them today.&lt;/li&gt;
&lt;li&gt;Reflect and connect, gain self-respect.&lt;/li&gt;
&lt;li&gt;Life moves fast? Make memories last.&lt;/li&gt;
&lt;li&gt;Get in the flow, write and grow.&lt;/li&gt;
&lt;li&gt;Pen in hand, understand.&lt;/li&gt;
&lt;li&gt;Daily lines, clear designs.&lt;/li&gt;
&lt;li&gt;Just take note, stay afloat.&lt;/li&gt;
&lt;li&gt;Write your heart, make a fresh start.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://x.com/i/grok/share/tmAxa8bTOl0kEcau3jqmLDEPT&#34;&gt;Grok 3&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;Write your day, let thoughts play!&lt;/li&gt;
&lt;li&gt;Pen in hand, your mind expands!&lt;/li&gt;
&lt;li&gt;Journal your way, seize the day!&lt;/li&gt;
&lt;li&gt;Feel the groove, let words move!&lt;/li&gt;
&lt;li&gt;Scribe your soul, make it whole!&lt;/li&gt;
&lt;li&gt;Ink flows free, just be me!&lt;/li&gt;
&lt;li&gt;Page by page, unlock your sage!&lt;/li&gt;
&lt;li&gt;Write it down, wear the crown!&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;At first glance, GPT 4.5 didn’t impress me. Claude 3.7 Sonnet did. I also didn’t like Gemini 2.5 Pro, but Grok was great.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Grok 3 &amp;gt; Claude 3.7 Sonnet &amp;gt; Gemini 2.5 Pro &amp;gt; GPT 4.5.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;But it’s hard to compare a dozen quotes at once. So I made a small &lt;a href=&#34;https://tools.s-anand.net/quotesarena/&#34;&gt;quotes arena app&lt;/a&gt; to help me pick my favorites. It shows me random pairs of quotes and asks which I like more.&lt;/p&gt;
&lt;p&gt;To my surprise, after answering 30+ &amp;ldquo;games&amp;rdquo; in the arena, I found that based on my preferences:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Claude 3.7 Sonnet &amp;gt; Gemini 2.5 Pro &amp;gt; GPT 4.5 &amp;gt; Grok 3.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That was weird. I thought I liked Grok&amp;rsquo;s results a lot. I continued till I answered 50+ games. Then I found that:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-text&#34; data-lang=&#34;text&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;Grok 3 &amp;gt; GPT 4.5 &amp;gt; Gemini 2.5 Pro &amp;gt; Claude 3.7 Sonnet.
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;That&amp;rsquo;s the &lt;strong&gt;exact&lt;/strong&gt; opposite of the previous result.&lt;/p&gt;
&lt;p&gt;Honestly, I&amp;rsquo;m depressed. I&amp;rsquo;ve learnt 3 things:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;I can&amp;rsquo;t judge stuff at a glance.&lt;/li&gt;
&lt;li&gt;But I think I can (especially with code.)&lt;/li&gt;
&lt;li&gt;Even when evaluating carefully, my preferences are unstable.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;strong&gt;Nothing&lt;/strong&gt; has shaken my confidence more in recent times. I &lt;strong&gt;cannot&lt;/strong&gt; trust my judgement. I need written evals. Badly.&lt;/p&gt;
</description>
    </item>
  </channel>
</rss>
