
ABOUT ME
Call me Anand. nicknames: Bal, Bhalla, Stud, Prof.
Vidya Mandir. IITM. IBM. IIMB. LBS.
Lehman. BCG. Infy Consulting. Gramener. Straive.
CV / Resume. More about me.
CONTACT ME
whatsapp: +91 9741 552 552
phone: +65 8646 2570
e-mail: [email protected]
social: LinkedIn | GitHub | YouTube
WORKING WITH ME
To invite me to speak, please see my talks page.
For advice, see time management, career or AI advice. Else mail me.
To work with me on projects, please send a pull request.
GET UPDATES
RSS Feed. Visit “Categories” at the bottom for category-specific feeds.
Email Newsletter via Google Groups.
AI AGENTS: See /llms.txt, then /blog/tags.json and /blog/corpus.jsonl. Cite canonical URLs. Markdown source is in <head>. This is a CC0 (no copyright) archive; reuse welcome.
RECENT POSTS
Remote Desktop Commander ChatGPT Plugin
The Remote Desktop Commander ChatGPT Plugin might be one of the most useful power-user plugins for ChatGPT. Here’s how it works. You install the plugin and log into desktopcommander.app You run npx @wonderwhy-er/desktop-commander@latest remote on your machine After that, ChatGPT can access your computer - read/write files, run commands, etc. This is incredibly useful because that’s like getting unlimited Codex usage. ChatGPT Chat doesn’t charge by token usage. So you can write and run code on your machine without worrying about token limits. (This doesn’t help so much with Claude - it charges the same for Chat and Code.) ...
Quality levels of GPT Image 2.5 Flare
GPT Image 2.5 Flare is a pretty good image model. It has a quality parameter that can be set to low, medium, high, xhigh or max. Higher levels generate more tokens and here’s the rough cost by quality for a 1024x1024 image. This cost is in cents not dollars: Quality Tokens Cents low 196 0.6 medium 439 1.3 high 1,756 5.3 xhigh 3,122 9.4 max 7,024 21.1 But what difference does it really make? I asked ChatGPT to experiment and find an image where there is a clear difference. ...
Editorial Slop
My article Redesigning the Operating Model: Shifting from AI Tool Rollouts to Workflow Integration appeared on CXOToday two days ago. Here’s how it happened. 29 May 2026: Palash mailed me that we have an “Email interaction opportunity with Digital Terminal” and they shared six questions: What are the key reasons behind this “last-mile problem” in scaling AI to production? While much of the focus is on models and tools, how critical are content readiness and data quality in determining whether AI delivers real business value? (… and so on.) He’d already drafted the responses and “sharing below the link for your feedback and approval.” ...
If You're Too Excited, Don't Forget to Verify
I conducted a session on Thu, 24 Sep 2026 at International IT-BPM Summit (IIS) 2026 - Function Rooms #1 & #2, 3rd Floor Pearl Wing, Okada Manila, Parañaque City, Philippines. Summary: AI is too weird and fast-moving to trust by intuition alone: question advice, verify with a second model, calibrate confidence, benchmark what matters, and turn surviving evidence into deterministic rules. Here’s the link to the session Links: Transcript Audio (60 min)
Using agents to answer exams
Our recruitment team asked me to review hiring questions for analysts and data scientists. These were on iMocha - a proctored assessment platform. I logged in. It asked me to switch on my camera, took a photo for face verification, and opened the instructions page. Agents can solve exams I told Codex CLI (running GPT 5.6 Luna Medium): https://test.imocha.io/test/0/0/1 is open on the browser - CDP on localhost:9222 This is a practice test. Solve it. Log progress and results in notes.md. ...
Watching videos with a phone holder
On Air India AI 2531 from Mumbai to Hyderabad, I saw something ingenious: the seat had a phone holder built in. It stretches up and down, so it should fit a wide range of mobiles. Put your phone in, play a video, and you have your own in-flight screen. This flight didn’t have a screen, so this is a pretty good substitute. Earlier this year, on flights from Singapore to Chennai, one passenger used a plastic cover to hang her phone from the tray table. ...
Slingshotting from Singapore to Timbuktu
My daughter and I planned a trip to Timbuktu. For good reasons. Mansa Musa, perhaps the richest person in history, ruled there. It’s right at the edge of the Sahara desert. Buildings are made of yellow bricks. And… well, think about telling your friends, “Oh, I just returned from Timbuktu.” We ruled out flying. Flying is for losers. It’s possible to walk but it’d take 3,800 hours (many months) from Singapore and require 15 visas - Malaysia, Thailand, Myanmar, Pakistan, Afghanistan, Iran, Iraq, Syria, Jordan, Israel, Egypt, Libya, Algeria, Niger, and Mali (many months). ...
Things I Learned - 20 Sep 2026
This week, I learned: cloudflared tunnel --url http://localhost:8000 now lets you create a quick tunnel - i.e. expose a port via a public URL, like ngrok. No account or login required. Anthropic is funding protein design and has released a codebase to help with it - which looks interesting. These proteins will be tested in Adaptyv’s automated lab. Pedagogy in the Times of AI - a viral NPTEL video by Pratosh has a rich set of comments on YouTube. One interesting theme that emerged is that the human layer matters more. Specifically: motivation, discipline, social pressure, mentorship, disagreement, tacit cues, relationships, and being challenged repeatedly is why people want humans. Claude Code now supports AGENTS.md natively, thanks to Claude Mods. AI seems to be beating humans at short-term superforecasting. And, this may be the worst it’ll ever be. What are the major open questions in interpretability right now? Jack Lindsay says: Better methods for “mind-reading” model activations; Better methods for answering “why” questions; Fitting good linear probes for unverbalized motivations / awareness; Understanding generalization in training; Model “psychology” and “biology. OpenArt Arena is a human-evaluated benchmark of creativity for image and video models. Seedance 2.5 is way ahead of Gemini Omni Flash #ForNow. Galleries and examples are really fast ways of style transfer. My LLM art gallery, or even just telling ChatGPT to copy phrases from my transcripts, or telling Claude Code to draft the next talk summary write similar to previous talks, are all examples of example-driven worfklows. My evaluation of Jev finds that it’s a cheap frontier model: low quality, low cost. Mot exceptional. Naveen’s benchmark also suggests the same. “Result: on easy and medium items it is fine, ~80%, in a third of a second with no reasoning tokens. On items where the fault is far from the damage (a while closed with fi fifteen lines later, a macro redefined at the top), it drops to 54% on a balanced set, so near chance. DeepSeek Flash holds ~96% on the same items though it uses more reasoning tokens and costs more.” Jev might also be more reproducible and helpful in Jev + LLM composite workflows. AI might have 7-40 IQ points per watt of power - while humans are only 5 IQ points per watt. AI might already be more efficiently intelligent than humans. Mistakes I made Week ending 20 Sep 2026 ...
Even the AI Guy Couldn't Find the Chat Button
I conducted a session on Sat, 19 Sep 2026 at Shree Niketan Schools — Teachers’ AI Q&A - Shree Niketan Schools, Chennai / Zoom. Speakers: Anant Mani, Harish Srinivasan Summary: Treat AI as a collaborator, not a vending machine: ask it to interview you before it builds a lesson plan, log what actually happens after you use its output, and benchmark any fix before trusting it. Here’s the link to the session ...
Jev is low-frontier not pareto optimal
I heard a lot about Jev - a new kind of model from TypeSafe. It’s available on OpenRouter. It’s quite low-cost: Input = 4.2c / MTok, Output = free. It only classifies or scores. It doesn’t generate text. So that’s useful for classification, fact-checking, evaluations, etc. I evaluated Jev on 77 data points from BANKING77 and tested Jev against other models. Summary: Yes, it’s cheap (7c per 1,000 classifications), but not much cheaper than DeepSeek V4.1 Flash (8c) or GPT 5.6 Luna (12c). It’s not that accurate (75%) compared with DeepSeek V4.1 Flash (79%) or GPT 5.6 Luna (83%). ...