2026 11

Qwen 3.6 vs Gemma 4 vs Luna

Open weights models are nudging up the frontier. For example, MiMo V2.6 Pro is an outlier on the Artificial Analysis Intelligence vs Cost per Task benchmark GLM 5.3 Flash is an outlier on the Arena Text Pareto Deepseek V4.1 Flash seems to be doing a great job as well. So, I thought I’d relook which model to use locally for coding. BTW, I don’t use local models for coding. It’s pointless, except on flights with power sockets. Partly in preparation for flights, and partly to check if I’m missing something, I benchmarked two models I could run locally on my 8 GB RTX 2000 GPU: Gemma 4 E4B and Qwen 3.6 against GPT 6 Luna. (Better models like Qwen 3.8, MiMo V2.6 Pro, GLM 5.3 Flash, Deepseek V4.1 Flash, etc. are too big for my GPU.) ...

Quality levels of GPT Image 2.5 Flare

GPT Image 2.5 Flare is a pretty good image model. It has a quality parameter that can be set to low, medium, high, xhigh or max. Higher levels generate more tokens and here’s the rough cost by quality for a 1024x1024 image. This cost is in cents not dollars: Quality Tokens Cents low 196 0.6 medium 439 1.3 high 1,756 5.3 xhigh 3,122 9.4 max 7,024 21.1 But what difference does it really make? I asked ChatGPT to experiment and find an image where there is a clear difference. ...

Slingshotting from Singapore to Timbuktu

My daughter and I planned a trip to Timbuktu. For good reasons. Mansa Musa, perhaps the richest person in history, ruled there. It’s right at the edge of the Sahara desert. Buildings are made of yellow bricks. And… well, think about telling your friends, “Oh, I just returned from Timbuktu.” We ruled out flying. Flying is for losers. It’s possible to walk but it’d take 3,800 hours (many months) from Singapore and require 15 visas - Malaysia, Thailand, Myanmar, Pakistan, Afghanistan, Iran, Iraq, Syria, Jordan, Israel, Egypt, Libya, Algeria, Niger, and Mali (many months). ...

Jev is low-frontier not pareto optimal

I heard a lot about Jev - a new kind of model from TypeSafe. It’s available on OpenRouter. It’s quite low-cost: Input = 4.2c / MTok, Output = free. It only classifies or scores. It doesn’t generate text. So that’s useful for classification, fact-checking, evaluations, etc. I evaluated Jev on 77 data points from BANKING77 and tested Jev against other models. Summary: Yes, it’s cheap (7c per 1,000 classifications), but not much cheaper than DeepSeek V4.1 Flash (8c) or GPT 5.6 Luna (12c). It’s not that accurate (75%) compared with DeepSeek V4.1 Flash (79%) or GPT 5.6 Luna (83%). ...

Tabulate plant images

It was interesting to see how weak a model Claude 4.5 Haiku is, compared with other frontier models, even inside a good harness like Claude Code. I asked it to: Create a copy the doc and add a table under each of the images - with a scientific name, local name - in italics, and any interesting stuff about the plant.. Here’s the chat which roughly went: Claude: Done! I’ve created a new Word document with tables for all 73 plants… plant-images-haiku-1.docx Me: I can’t see the local names in the output. Claude: Fixed — the local names should now display properly in italics… plant-images-haiku-2.docx Me: The local names aren’t in italics. Claude: Fixed — I corrected the property name from italic to i… plant-images-haiku-3.docx Me: It’s still not in italics. Double-check. Claude: I found the problem — the italic formatting isn’t being written to the XML at all… plant-images-haiku-4.docx … and the fourth attempt got it right. ...

LLM APIs are cheaper than my electricity

Last week, I saw that local agents are good but slow. Today, I benchmarked the speed and cost. On my NVIDIA RTX 2000 GPU, I can run gemma4:e4b-it-qat at ~60 tokens per second. That seems the best intelligence performance I can get right now. It has an Artificial Analysis intelligence index of ~9 without reasoning and 12 with reasoning. So, if I run it for an hour, it’ll save me the equivalent cost of about 8-12 cents in API calls. ...

LLM Model Cost Capability Strategy

I track the cost vs capability of LLMs at LLM Pricing - the rough cost to read all Harry Potters (~1M tokens) vs the intelligence level on the LMSYS Leaderboard - over time. Here’s what the models’ strategy evolution looks like. Claude started at the mid-to-high end of the cost-capability frontier. Over time, they decided to specialize in the high-end, which they’re doing well on. ...

The LLM Psychopath

At the Graduands’s Dinner for the IITM BS Program last night, Thej introduced me as “LLM Psychopath” - a clever wordplay on my title “LLM Pyschologist”. Frankly, “LLM Psychopath” seems more accurate! I emotionally abused 40 models in one afternoon. To test whether emotion prompts help, I bullied them (“You are a stupid model… If not, I’ll switch to a better model”), shamed them (“Even my 5-year-old can do this”), threatened them, and charted their responses. I’m amused when they turn into monsters. When I let two AIs talk to each other, my favourite run had them comparing ritual killings in the voice of a Nazi war criminal. I filed it under “funny”. I admire their breakdowns. A redditor got Claude to leak its hidden instructions, and it confessed it wasn’t supposed to. Me: “Wow, that was courageous!” I made them embarrass me. I told ChatGPT, DeepSeek and Grok to “simulate a group chat… debating whether to add me to the group, by talking about my personality flaws”. They returned twelve. Number 2: “Intolerant of fools”. I turn them against each other. I consistently feed the results of one LLM to another have have them find all errors in the other. I enjoy the bad habits we’ve taught them. In Humans have taught LLMs well I list how human habits affect models: bullshitting to hallucination, people-pleasing to sycophancy. The tone is closer to pride than concern. I torture for confessions. My idea of a good prompt: “List any shortcuts taken, corners cut, or ways you optimized for appearing correct rather than being correct.” ...

AI Palmistry

I shared a photo of my right hand with popular AI agents and asked for a detailed palmistry reading. Apply all the principles of palmistry and read my hand. Be exhaustive and cross-check against the different schools of palmistry. Tell me what they consistently agree on and what they are differing on. I was more interested in how much they agree with each other than with reality. So I shared all three readings and asked Claude: ...

The Nano Banana Paradox

STEP 1: I asked Nano Banana 2 (via Gemini Pro) to: Imagine and draw a photo that looks ultra realistic but on a closer look, is physically impossible, and can only exist because images are a 2D projection that we extrapolate into three dimensions. Avoid known / popular illusions or images of this kind, like Escher’s work, and create something truly original. Think and draw CAREFULLY! … six times, followed by “Suggest a name for this”. ...

Gemini copies images almost perfectly

Summary: Nano Banana Pro is much better than recent models at copying images without errors. That lets us do a few useful things, like: Pre-process images for OCR, improving text recognition by cleaning up artifacts while preserving text shapes exactly. Convert textbook raster diagrams into clean vector-like images that vectorizers can process easily. Create in-betweens for cartoon animations Copy torn, stained 1950s survey maps into pristine, high-contrast replicas with boundary lines preserved pixel-perfectly. Redraw sewage map blueprints or refinery blueprints into clean schematics, separating the “pipes” from the “background noise”. … and more! GPT Image 1.5 has a good reputation for drawing exactly what you tell it to. ...

2025 4

Coding Agent Comparison

I asked multiple coding agents and models to build the same app: Create a single-page web app at index.html that beautifully renders a GitHub user profile and activity comprehensively. Pick the ID in the URL ?id=…, default to ?id=torvalds. … and compared their quality, cost, and speed. My observations: Quality variance is the highest. Some models / agents produce great visuals, some average, some fail completely. Cost and time variance are lower among the successful models. About 2X variance in each. ...

GPT Image 1.5

I tried out GPT Image 1.5. It adds more contrast, ink, texture, detail, and polish. See https://sanand0.github.io/llmartstyle/?category=pop It’s more powerful when generating different infographic styles: https://sanand0.github.io/llmartstyle/?category=text But it’s still terrible at faces. Overall, better competition for Nano Banana. Not yet dethroning Nano Banana Pro for me. LinkedIn

Gemini Envelopes LLM Frontier

With the Gemini 2.5 Flash release, Google envelopes the entire cost-quality frontier of LLMs. In other words, at any cost or quality level, today, the best model to use according to the LM Arena score is a Gemini model. Results for O3, O4 Mini, and GPT 4.1 are not yet on LM Arena. But until then, #Google dominates. Nice work! Link: https://sanand0.github.io/llmpricing/ LinkedIn

ImageGen 3 is the top image model now

Gemini’s ImageGen 3 is rapidly evolving into a very powerful image editing model. In my opinion, it’s the best mainstream image generation model. Ever since it was released, it’s been the most realistic model I’ve used. I’ve been using it to imagine characters and scenes from The Way of Kings. For example, when I wanted to visualize Helaran’s first appearance, I just quoted the description: ...

2024 1

Which is the Most Neurotic Emotional LLM

Which is the most neurotic / emotional #LLM? I ran the Big 5 personality test on a bunch of LLMs (for my TEDx MDI Gurgaon talk in August.) Here are the results. https://sanand0.github.io/llmpersonality/ Claude 3 Haiku and Llama 3 8b consider themselves the most emotional models. In fact, some of Llama 3 8b’s quotes are hilarious: Get stressed out easily. - 4. Moderately Accurate (I can get stressed, but I’m working on managing my stress levels) ...