<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>gemini-cli on S Anand</title>
    <link>https://www.s-anand.net/blog/tag/gemini-cli/</link>
    <description>Recent content in gemini-cli on S Anand</description>
    <generator>Hugo -- 0.164.0</generator>
    <language>en-us</language>
    <lastBuildDate>Sun, 08 Mar 2026 15:09:25 +0800</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/tag/gemini-cli/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Gemini CLI harness is not good enough</title>
      <link>https://www.s-anand.net/blog/gemini-cli-harness-is-not-good-enough/</link>
      <pubDate>Sun, 08 Mar 2026 15:09:25 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/gemini-cli-harness-is-not-good-enough/</guid>
      <description>&lt;p&gt;I&amp;rsquo;ve long felt that while the Gemini 3 Pro model is fairly good, the Gemini CLI harness isn&amp;rsquo;t. I saw an example of this today.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Me&lt;/strong&gt;: Tell me the GitHub IDs of all students in this directory.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gemini CLI&lt;/strong&gt;:&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;SearchText &amp;#39;github&amp;#39; within ./
Found 100 matches (limited)
Sending this message (14606686 tokens) might exceed the remaining context window limit (1037604 tokens).
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;&lt;strong&gt;Me&lt;/strong&gt;: Only send the (small) required snippets of data. Write code as required.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gemini CLI&lt;/strong&gt;:&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;SearchText &amp;#39;github&amp;#39; within ./
Found 100 matches (limited)
Sending this message (14606686 tokens) might exceed the remaining context window limit (1037604 tokens).
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://files.s-anand.net/images/2026-03-08-gemini-cli-harness.avif&#34;&gt; &lt;!-- https://gemini.google.com/app/40182d961e78af0d --&gt;&lt;/p&gt;
&lt;p&gt;Come ON! It&amp;rsquo;s &lt;strong&gt;March 2026&lt;/strong&gt;. We can&amp;rsquo;t pretend it&amp;rsquo;s October 2025 any more.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;PS: I partly take back what I said. Codex had trouble, too. This problem may be harder than I thought. Still, Gemini CLI should not have gotten stuck where it did.&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 09 Nov 2025</title>
      <link>https://www.s-anand.net/blog/things-i-learned-09-nov-2025/</link>
      <pubDate>Sun, 09 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-09-nov-2025/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;But when an identity based belief was challenged, the brain responded as if under physical attack.&amp;rdquo; &lt;a href=&#34;https://spf13.com/p/the-hidden-conversation/&#34;&gt;Why Engineers Can&amp;rsquo;t Be Rational About Programming Languages&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Notes from &lt;a href=&#34;https://www.youtube.com/watch?v=I9Njb8Lw5Xc&#34;&gt;How to build a cult&lt;/a&gt;, Lulu Cheng, The Knowledge Project podcast
&lt;ul&gt;
&lt;li&gt;Conviction is infectious.&lt;/li&gt;
&lt;li&gt;Communicate at the INTERSECTION of interests. Learn theirs&lt;/li&gt;
&lt;li&gt;Begin with &amp;ldquo;why your story matters to them&amp;rdquo; (first sentence). That beats &amp;ldquo;how you tell it&amp;rdquo; &amp;gt; &amp;ldquo;where you tell it&amp;rdquo;.&lt;/li&gt;
&lt;li&gt;The easiest way to align with an audience is to find your community.&lt;/li&gt;
&lt;li&gt;Humor, curiosity, awe, any strong emotion is a hook.&lt;/li&gt;
&lt;li&gt;Culture has momentum. Best way to break it is to show an alternative that works. People will copy that&lt;/li&gt;
&lt;li&gt;REPEAT messages over and over with complete CONVICTION to convince people who TRUST you. That works, but you need all three.&lt;/li&gt;
&lt;li&gt;Trust builds from likeability, repeated exposure, common beliefs.&lt;/li&gt;
&lt;li&gt;An excellent way to defend against online criticism (when it matters) is to just SHOW UP and THANK them for feedback.&lt;/li&gt;
&lt;li&gt;Serious reputational damage must either be fixed immediately - or you live with it forever.&lt;/li&gt;
&lt;li&gt;Between a story and statistics, the story will always wins. Never fight a story with a statistic. Dig into your statistics and uncover BETTER stories.&lt;/li&gt;
&lt;li&gt;⭐ Prebuttals are a great idea. Start with all possible criticisms yourself and diffuse them. The other person has nothing left to say&lt;/li&gt;
&lt;li&gt;Sparring keeps you sharp. Spar with LLMs.&lt;/li&gt;
&lt;li&gt;To defend, show how the attack targets other people, increasing the surface area. Show how the SPECIFIC attack targets a larger group. Create a SPECIFIC cause worth fighting for.&lt;/li&gt;
&lt;li&gt;Each role has specific objective to optimise for. The leader&amp;rsquo;s role is to balance across these.&lt;/li&gt;
&lt;li&gt;Cheerleader effect. People look beautiful next to a cheerleader. Associations taint.&lt;/li&gt;
&lt;li&gt;Each person has dozens of aspects to their persona. We cannot remember all of them. Each person can make a choice on who they project themselves to be in any group. Shaping their persona.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;del&gt;The &lt;a href=&#34;https://marketplace.visualstudio.com/items?itemName=mechatroner.rainbow-csv&#34;&gt;Rainbow CSV&lt;/a&gt; extension may be causing delays (infinite spinner) when pasting Markdown in VS Code. Restarting it seems to fix the issue.&lt;/del&gt;&lt;/li&gt;
&lt;li&gt;⭐ &lt;a href=&#34;https://github.com/K-Dense-AI/claude-scientific-skills/&#34;&gt;Claude scientific skills&lt;/a&gt; is a collection of skills teaching Claude how to use scientific libraries, databases, and APIs across several domains. This may be a good example of a non-trivial skill library - that is hard for AI coding agents to infer by themselves.&lt;/li&gt;
&lt;li&gt;Notes from &lt;a href=&#34;https://blog.sshh.io/p/how-i-use-every-claude-code-feature&#34;&gt;How I use every Claude Code feature&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Use AGENTS.md as guardrails, not a manual. Document what it gets wrong.&lt;/li&gt;
&lt;li&gt;Use self-documenting tools/APIs rather than documenting.&lt;/li&gt;
&lt;li&gt;Docs: Explain why and when to read each doc.&lt;/li&gt;
&lt;li&gt;Never say &amp;ldquo;Never.&amp;rdquo; Explain when to which which alternative.&lt;/li&gt;
&lt;li&gt;Prefer CLIs for stateless tools, MCPs for stateful, authenticated, or complex (e.g. Playwright).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Coding agents work well with version control. &lt;a href=&#34;https://simonwillison.net/2025/Nov/4/datasette-10a20/&#34;&gt;Simon Willison&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Break up uncommitted changes into small commits&lt;/li&gt;
&lt;li&gt;Rewrite branch history for readability&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;gh&lt;/code&gt; CLI to fetch line-wise comments from a PR and make requested changes (e.g. renaming, refactoring, adding types, etc.)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;⭐ When using MCPs or tools with private data, &amp;ldquo;color untrusted content in red, unsafe actions in blue, and never mix colors.&amp;rdquo; &lt;a href=&#34;https://timkellogg.me/blog/2025/11/03/colors&#34;&gt;Good advice&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;⭐ DeepWiki offers a &lt;a href=&#34;https://cognition.ai/blog/codemaps&#34;&gt;codemaps&lt;/a&gt; feature that explains code in an interactive way. It shows a structured explanation on the left. You can click on any note to see the code on the right. It&amp;rsquo;s an effective way to understand how a library or tool executes a task. &lt;a href=&#34;https://deepwiki.com/search/draw-a-codemap_59d591f6-fc79-40f0-973d-bfa0e149b41a&#34;&gt;Here&amp;rsquo;s an example of how Mermaid works&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://blog.google/technology/developers/file-search-gemini-api/&#34;&gt;Gemini offers RAG with free storage&lt;/a&gt;. RAG costs are quite high. This simplifies the process a lot. But I tried running the sample program and after an hour, it still had not completed uploading a single file. Best to wait and watch.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/models?fmt=cards&amp;amp;order=top-weekly&amp;amp;output_modalities=embeddings&#34;&gt;OpenRouter supports embedding models&lt;/a&gt; using an &lt;a href=&#34;https://openrouter.ai/docs/api-reference/embeddings/create-embeddings&#34;&gt;OpenAI-like API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://openrouter.ai/moonshotai/kimi-k2-thinking&#34;&gt;Kimi K2 Thinking&lt;/a&gt; seems popular because
&lt;ul&gt;
&lt;li&gt;It&amp;rsquo;s an open-weights model on par with the top models on Humanity’s Last Exam (text-only) and BrowseComp&lt;/li&gt;
&lt;li&gt;Can run 200-300 tool calls without human guidance&lt;/li&gt;
&lt;li&gt;4x cheaper than GPT-5 with low tokens (32B active on 1T parameters, INT4 quantized)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Based on responses to &lt;a href=&#34;https://x.com/simonw/status/1979254349235925084&#34;&gt;Simon Willison&amp;rsquo;s question&lt;/a&gt;, &lt;a href=&#34;https://chatgpt.com/share/690b4fa0-7dec-800c-87a5-6e01dda36f7e&#34;&gt;ChatGPT&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Fine-tuning helps when:
&lt;ul&gt;
&lt;li&gt;Lower latency, e.g. for type-ahead, at lower cost (&lt;strong&gt;37 mentions&lt;/strong&gt;)&lt;/li&gt;
&lt;li&gt;Structured extraction, parsing and classifiers, e.g. postal address, detecting secrets (18 mentions)&lt;/li&gt;
&lt;li&gt;Custom vision models, e.g. check containers (12 mentions)&lt;/li&gt;
&lt;li&gt;Domain-specific code and stacks (niche languages, stack-specific generation, text→SQL) (11 mentions)&lt;/li&gt;
&lt;li&gt;&amp;hellip; and a long tail.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Fine tuning does not help:
&lt;ul&gt;
&lt;li&gt;When A base model plus prompting or RAG does as well or better (&lt;strong&gt;15 mentions&lt;/strong&gt;)&lt;/li&gt;
&lt;li&gt;When you risk being leapfrogged by a new release (4 mentions)&lt;/li&gt;
&lt;li&gt;When cost and data do not justify the ROI (3 mentions)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The data I can export from my Android phone includes the below. 🟢 indicates it&amp;rsquo;s tracked. 🟡 might need action, e.g. enabling / coding. &lt;a href=&#34;https://chatgpt.com/c/69089221-9430-8320-9cb0-5350a17fc486&#34;&gt;#&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;🟢 GPS/GNSS location (current &amp;amp; history). Turn on device Location. If you want a timeline you can export, enable Google Location History and later export via Google Takeout → Location History (JSON/KML).&lt;/li&gt;
&lt;li&gt;🟡 GNSS raw measurements (engineering traces). Android exposes GNSS “raw” logs on many devices; capture with dev tools or logging apps if supported (intended for research). See GNSS Raw Measurements API.&lt;/li&gt;
&lt;li&gt;🟢 Wi-Fi scans (nearby SSIDs/BSSIDs). Toggle Location scanning → Wi-Fi scanning in Location settings; apps need location permission to read results.&lt;/li&gt;
&lt;li&gt;🟡 Wi-Fi RTT distance to APs (indoor ranging). Apps can use Wi-Fi RTT (802.11mc/az) to measure distance to compatible APs; requires location permission.&lt;/li&gt;
&lt;li&gt;🟢 Bluetooth proximity/traffic. For packet-level logs, enable Developer options → Enable Bluetooth HCI snoop log, then pull &lt;code&gt;/sdcard/btsnoop_hci.log&lt;/code&gt; (Wireshark).&lt;/li&gt;
&lt;li&gt;🟢 Cell towers (IDs, signal strength). Apps can read via TelephonyManager (e.g., &lt;code&gt;getAllCellInfo()&lt;/code&gt;), with appropriate telephony permissions.&lt;/li&gt;
&lt;li&gt;🟢 Activity recognition (walking, running, in vehicle). Apps must request ACTIVITY_RECOGNITION (runtime) from Android 10+.&lt;/li&gt;
&lt;li&gt;🟢 Steps (step counter / detector). Use sensors API; from Android 10+ you must declare ACTIVITY_RECOGNITION to access step counter/step detector.&lt;/li&gt;
&lt;li&gt;🟢 Accelerometer / gyroscope / magnetometer streams. Apps read via SensorManager; some high-rate reads require HIGH_SAMPLING_RATE_SENSORS.&lt;/li&gt;
&lt;li&gt;🟢 Ambient light / proximity. Read via SensorManager; typically no special permission.&lt;/li&gt;
&lt;li&gt;🟢 Google Fit data (steps, workouts, heart rate from wearables, etc.). Manage and export from Google Fit / Google account Download your data.&lt;/li&gt;
&lt;li&gt;🟢 Contacts. MIUI → Settings → System apps → Contacts → Import/Export to .vcf (vCard).&lt;/li&gt;
&lt;li&gt;🟢 Call history / SMS (device). MIUI local/cloud backup can include call logs &amp;amp; messages; export by creating a local/Cloud backup and downloading. Note: 3P apps can’t read call/SMS logs unless they’re the default dialer/SMS.&lt;/li&gt;
&lt;li&gt;🟡 Gmail, Calendar, Contacts (Google). Export via Google Takeout (MBOX/ICS/CSV etc.).&lt;/li&gt;
&lt;li&gt;🟡 WhatsApp / Telegram / Signal chats. Use in-app exports: WhatsApp → Export chat, Telegram Desktop → Export, Signal → encrypted backup.&lt;/li&gt;
&lt;li&gt;🟢 Advertising ID. View/reset in Settings → Google → Ads (wording varies), per Google help on Ad ID reset.&lt;/li&gt;
&lt;li&gt;🟡 Per-app screen time / unlocks / opens. Third-party “usage” apps (e.g., analytics or “digital wellbeing” clones) require Usage Access (PACKAGE_USAGE_STATS). Use Android’s UsageStatsManager or apps that export CSV. Stock Digital Wellbeing does not offer an export.&lt;/li&gt;
&lt;li&gt;🟡 Notification history (last 24h). Settings → Notifications → Notification history → On. OEM-optional, but present on most devices. Viewable once enabled.&lt;/li&gt;
&lt;li&gt;🟡 Notification content stream (live). Grant an app Notification access to capture/export notifications going forward. (User-granted API via NotificationListenerService.) |&lt;/li&gt;
&lt;li&gt;🟢 Per-app data usage (mobile/Wi-Fi). Apps/ADB can query NetworkStatsManager; Settings shows per-app totals. Advanced dumps via &lt;code&gt;adb shell dumpsys netstats&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;🟡 Wi-Fi detailed logs. Developer options → Enable Wi-Fi verbose logging for richer diagnostics.&lt;/li&gt;
&lt;li&gt;🟡 Bluetooth packet logs. Developer options → Enable Bluetooth HCI snoop log; export file and analyze in Wireshark.&lt;/li&gt;
&lt;li&gt;🟢 Per-app storage usage. Apps/ADB can query StorageStatsManager; Settings shows per-app storage.&lt;/li&gt;
&lt;li&gt;🟡 Photo/video metadata (EXIF incl. location). Enable “Save location” in Camera app to embed GPS in EXIF; export files normally (EXIF remains). |&lt;/li&gt;
&lt;li&gt;🟢 Downloads &amp;amp; file metadata. Use a file manager or connect via USB; metadata is in the files themselves. |&lt;/li&gt;
&lt;li&gt;🟢 Battery usage history (per-UID/app), wakelocks, jobs. Generate adb bugreport and analyze with Battery Historian or &lt;code&gt;dumpsys batterystats&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;🟡 System/device logs (logcat). You can view via ADB/Android Studio. Android restricts 3rd-party access to system-wide logs for privacy.&lt;/li&gt;
&lt;li&gt;🟢 Developer quick tiles (Sensors off). Developer options → Quick settings developer tiles → Sensors off to globally cut Camera/Mic &amp;amp; SensorManager sensors on demand.&lt;/li&gt;
&lt;li&gt;🟡 Google Takeout: one-stop export for Location History (Timeline), Gmail (MBOX), Calendar (ICS), Google Photos, Drive, YouTube, Fit, etc.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://play.google.com/store/apps/details?id=com.arlosoft.macrodroid&#34;&gt;MacroDroid&lt;/a&gt;, &lt;a href=&#34;https://play.google.com/store/apps/details?id=com.llamalab.automate&#34;&gt;Automate&lt;/a&gt; and &lt;a href=&#34;https://play.google.com/store/apps/details?id=net.dinglisch.android.taskerm&#34;&gt;Tasker&lt;/a&gt; sound like powerful Android workflow automation tools. Some uses I can put it to:
&lt;ul&gt;
&lt;li&gt;Automatically upload recordings to Dropbox&lt;/li&gt;
&lt;li&gt;Turn off hotspot when I reach office&lt;/li&gt;
&lt;li&gt;Vibrate if I&amp;rsquo;m walking slowly&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Adding &lt;code&gt;&amp;lt;link rel=&amp;quot;alternate&amp;quot; type=&amp;quot;text/markdown&amp;quot; title=&amp;quot;LLM-friendly version&amp;quot; href=&amp;quot;/llms.txt&amp;quot;&amp;gt;&lt;/code&gt; is an emerging approach for pointing to LLMs.txt. It works. I asked Codex to read the &lt;a href=&#34;https://developers.cloudflare.com/workers/testing/vitest-integration/write-your-first-test/&#34;&gt;CloudFlare vitest page&lt;/a&gt;. It read the file truncating the middle, found the &lt;code&gt;&amp;lt;link rel=&amp;quot;alternate&amp;quot; type=&amp;quot;text/markdown&amp;quot; href=&amp;quot;https://developers.cloudflare.com/workers/testing/vitest-integration/write-your-first-test/index.md&amp;quot;/&lt;/code&gt; link in it, and reasoned &amp;ldquo;Considering fetching markdown instructions&amp;rdquo; and fetched the Markdown page. &lt;a href=&#34;https://www.gilesthomas.com/website-design&#34;&gt;Giles&amp;rsquo; Blog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/toon-format/toon&#34;&gt;toon&lt;/a&gt; is a YAML-like format that&amp;rsquo;s LLM friendly and especially token-efficient (CSV-like) for tables. You can convert back and forth between JSON and toon.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=2kCjSq_l-0s&#34;&gt;Food printing&lt;/a&gt; applies 3D printing techniques to create real food items. Given the art that this can create, I expect at least some adoption in niche restaurants.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/protomaps/PMTiles&#34;&gt;PMTiles&lt;/a&gt; lets you store map tiles as a single-file archive that libraries like MapLibre can read. Useful to avoid tile servers.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://mirrow.app/&#34;&gt;Mirrow&lt;/a&gt; is a CLI SVG animation builder that converts a DSL to animated SVGs. However, it may be easier to use an LLM to create the animated SVG directly with &lt;a href=&#34;https://developer.mozilla.org/en-US/docs/Web/SVG/Guides/SVG_animation_with_SMIL&#34;&gt;SMIL&lt;/a&gt; than learning Mirrow (or teaching the LLM Mirrow).&lt;/li&gt;
&lt;li&gt;⭐ One approach to giving memory (&amp;ldquo;episodic memory&amp;rdquo;) to coding agents is to &lt;a href=&#34;https://blog.fsck.com/2025/10/23/episodic-memory/&#34;&gt;allow them to search their logs&lt;/a&gt;.This gives them access to past discussions about a repo or other repos.&lt;/li&gt;
&lt;li&gt;To &lt;a href=&#34;https://github.com/google-gemini/gemini-cli/blob/main/docs/get-started/configuration.md&#34;&gt;configure Gemini CLI&lt;/a&gt; with an AI router, set:
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;&amp;quot;security.auth.selectedType&amp;quot;: &amp;quot;gemini-api-key&amp;quot;&lt;/code&gt; in &lt;code&gt;~/.gemini/settings.json&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;export GOOGLE_GEMINI_BASE_URL=https://llmfoundry.straive.com/gemini/&lt;/code&gt; (or your AI router base URL for Gemini)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;export GEMINI_API_KEY=...&lt;/code&gt; (your AI router API key)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Passing a HAR export to an LLM to build a scraper is a powerful idea! &lt;a href=&#34;https://youtu.be/7NCUE02l1DE?t=516&#34;&gt;Lessons from Diagram Chasing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Addy Osmani&amp;rsquo;s &lt;a href=&#34;https://github.com/addyosmani/gemini-cli-tips&#34;&gt;Gemini CLI tips&lt;/a&gt; are practical guides to using any coding agent, not just Gemini. I learnt about:
&lt;ul&gt;
&lt;li&gt;Run shell commands with &lt;code&gt;!&lt;/code&gt;, e.g. &lt;code&gt;!ls -la&lt;/code&gt; or even &lt;code&gt;!bash&lt;/code&gt;. It&amp;rsquo;s added to the chat.&lt;/li&gt;
&lt;li&gt;On-the-fly tool creation: ask it to write code for the task on the fly.&lt;/li&gt;
&lt;li&gt;Use it for system optimization, e.g. editing dotfiles, system customization, log error analysis, etc.&lt;/li&gt;
&lt;li&gt;Run &lt;code&gt;GEMINI_SYSTEM_MD=... gemini -p &amp;quot;task&amp;quot; --yolo --format json &amp;lt; input.txt&lt;/code&gt; to run Gemini with a different system prompt and feed it input.txt to run in a pipeline. (FYI: Codex does not send a default system prompt, so there&amp;rsquo;s nothing to override.)&lt;/li&gt;
&lt;li&gt;There is a &lt;a href=&#34;https://github.com/google-gemini/gemini-cli/discussions/categories/show-and-tell?discussions_q=is%3Aopen+category%3A%22Show+and+tell%22+sort%3Atop&#34;&gt;Gemini CLI Show and Tell&lt;/a&gt; thread with examples. This include &lt;a href=&#34;https://github.com/google-gemini/gemini-cli/discussions/7890&#34;&gt;Janitor AI&lt;/a&gt;, a &lt;a href=&#34;https://github.com/google-gemini/gemini-cli/discussions/3965&#34;&gt;Gemini CLI session viewer&lt;/a&gt;, etc.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://codelabs.developers.google.com/gemini-cli-hands-on&#34;&gt;Hands on with Gemini CLI&lt;/a&gt; has several &lt;a href=&#34;https://codelabs.developers.google.com/gemini-cli-hands-on#11&#34;&gt;Use cases to try out&lt;/a&gt;.
&lt;a href=&#34;https://github.com/amitkmaraj/gemini-cli-custom-slash-commands/blob/main/.gemini/commands/photo-rename.toml&#34;&gt;Renaming photos&lt;/a&gt; and
&lt;a href=&#34;https://github.com/amitkmaraj/gemini-cli-custom-slash-commands/blob/main/.gemini/commands/file-organizer.toml&#34;&gt;organizing files&lt;/a&gt; are clever ones.&lt;/li&gt;
&lt;li&gt;AGENTS.md can be used like a decision log - rules, styles, or preferences that evolve over time - on a per-repo basis. Gemini&amp;rsquo;s &lt;code&gt;/memory add&lt;/code&gt; feature helps with this.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;gemini --checkpointing&lt;/code&gt; is a useful &amp;ldquo;undo&amp;rdquo; feature. &lt;code&gt;/restore&lt;/code&gt; rolls you back to a specific checkpoint. The overhead is small.&lt;/li&gt;
&lt;li&gt;Caching is only available with API key or Vertex AI, not OAuth login &lt;a href=&#34;https://google-gemini.github.io/gemini-cli/docs/cli/token-caching.html&#34;&gt;as of now&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.s-anand.net/blog/openai-tts-cost/&#34;&gt;OpenAI TTS costs are confusing&lt;/a&gt;. But in short
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.openai.com/docs/models/tts-1&#34;&gt;TTS-1&lt;/a&gt; costs $15 / MChars (max 4,096 chars per request), which ends up at ~86c / hour&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.openai.com/docs/models/gpt-4o-mini-tts&#34;&gt;GPT-4o Mini TTS&lt;/a&gt; costs ~$16 / MChars (max 2K tokens which is ~7,000 chars per request), which ends up at ~88c / hour. Very similar cost, effectively&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.openai.com/docs/models/tts-1-hd&#34;&gt;TTS-1 HD&lt;/a&gt; is twice TTS-1.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI has a &lt;a href=&#34;https://platform.openai.com/docs/api-reference/usage/&#34;&gt;usage API&lt;/a&gt; that provides &lt;a href=&#34;https://platform.openai.com/docs/api-reference/usage/cost&#34;&gt;cost&lt;/a&gt; as well as usage for &lt;a href=&#34;https://platform.openai.com/docs/api-reference/usage/completions&#34;&gt;completions&lt;/a&gt;, &lt;a href=&#34;https://platform.openai.com/docs/api-reference/usage/images&#34;&gt;images&lt;/a&gt;, &lt;a href=&#34;https://platform.openai.com/docs/api-reference/usage/audio_speeches&#34;&gt;audio speeches&lt;/a&gt;, etc.
&lt;ul&gt;
&lt;li&gt;These require an &lt;a href=&#34;https://platform.openai.com/settings/organization/admin-keys&#34;&gt;organization admin key&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Cost API: &lt;code&gt;curl &amp;quot;https://api.openai.com/v1/organization/costs?start_time=$TIMESTAMP&amp;amp;project_ids=$PROJECT_ID&amp;amp;group_by=line_item&amp;quot;&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Audio speech usage API: &lt;code&gt;curl &amp;quot;https://api.openai.com/v1/organization/usage/audio_speeches?start_time=$TIMESTAMP&amp;amp;project_ids=$PROJECT_ID&amp;amp;group_by=model&amp;quot;&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 03 Aug 2025</title>
      <link>https://www.s-anand.net/blog/things-i-learned-03-aug-2025/</link>
      <pubDate>Sun, 03 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-03-aug-2025/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;From &lt;a href=&#34;https://www.newyorker.com/magazine/2025/07/21/ai-is-about-to-solve-loneliness-thats-a-problem&#34;&gt;A.I. Is About to Solve Loneliness. That’s a Problem&lt;/a&gt;: “Blindly stifling every flicker of boredom with enjoyable but empty distractions precludes deeper engagement with the messages boredom sends us about meaning, values, and goals.” Maybe the best thing about boredom is what it forces us to do next.&lt;/li&gt;
&lt;li&gt;Here&amp;rsquo;s when be candid vs polite. #beliefs &lt;a href=&#34;https://chatgpt.com/share/688e29be-d4bc-800c-b5f5-527c3502bf78&#34;&gt;ChatGPT&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;If there&amp;rsquo;s high trust (i.e. the other person trusts you):
&lt;ul&gt;
&lt;li&gt;Important topic/decision: Be candid&lt;/li&gt;
&lt;li&gt;Unimportant: Follow culture (e.g. in Japan, you&amp;rsquo;d be polite; in The Netherlands, you&amp;rsquo;d be candid)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Low trust:
&lt;ul&gt;
&lt;li&gt;Important: Earn trust first&lt;/li&gt;
&lt;li&gt;Unimportant: Be polite&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;I didn&amp;rsquo;t realize that it was &lt;a href=&#34;https://en.wikipedia.org/wiki/Luis_Walter_Alvarez&#34;&gt;Luis Alvarez&lt;/a&gt; (whom I know from his work on the bubble chamber) is the &lt;em&gt;same&lt;/em&gt; person who figured out that &lt;a href=&#34;https://en.wikipedia.org/wiki/Alvarez_hypothesis&#34;&gt;an asteroid killed dinosaurs&lt;/a&gt;. He also used muon tomography to search pyramids for hidden chambers and figured out Kennedy was shot from behind. Added his biography, &lt;a href=&#34;https://www.goodreads.com/book/show/218569821-collisions&#34;&gt;Collisions&lt;/a&gt; to my &lt;a href=&#34;https://www.goodreads.com/review/list/39713492-s-anand?ref=nav_mybooks&amp;amp;shelf=to-read&amp;amp;sort=date_added&#34;&gt;to-read list&lt;/a&gt;. &lt;a href=&#34;https://en.wikipedia.org/wiki/Luis_Walter_Alvarez#Scientific_detective_work&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Benjamin Green &lt;a href=&#34;https://resobscura.substack.com/p/openais-new-study-mode-and-the-risks&#34;&gt;suggests&lt;/a&gt; that &lt;a href=&#34;https://openai.com/index/chatgpt-study-mode/&#34;&gt;OpenAI Study mode&lt;/a&gt; is sycophantic. E.g. in &lt;a href=&#34;https://chatgpt.com/share/688a9730-85d0-8004-9dae-0edb0c3ceff4&#34;&gt;this conversation&lt;/a&gt;, ChatGPT &lt;em&gt;carefully&lt;/em&gt; balances truth and politeness. A reader might misinterpret that as agreement. But sometimes, we &lt;em&gt;need&lt;/em&gt; candor. Politeness trades clarity for harmony. &lt;strong&gt;People who trust AI should tell it to be more candid&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;⭐ Here&amp;rsquo;s my current response when asked, &amp;ldquo;How should I use LLMs better&amp;rdquo;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Use the best models, consciously&lt;/strong&gt;. O3 (via $20 ChatGPT), Gemini 2.5 Pro (free on Gemini app), or Claude 4 Opus (via $20 Claude). The older models are the default and far worse.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Speak &amp;amp; listen, don&amp;rsquo;t just type &amp;amp; read&lt;/strong&gt;. I had to resist the temptation to ignore ChatGPT response when a colleague read it out. We are patient with and have respect for humans but not for AI. The value we derive requires both. Suggestion: Speak and listen rather than type and read. It&amp;rsquo;s hard to skip and easier to stay in the present. It&amp;rsquo;s also easier to ramble than type.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Keep an impossibility list&lt;/strong&gt;. There is a jagged edge that moves. When you note down what&amp;rsquo;s impossibile today and retry every month, you can see how that edge shifts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Wait for better models&lt;/strong&gt;. Many problems can be solved just by waiting a few months for a new model. You don&amp;rsquo;t need to find or build your own app.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Make context easily available&lt;/strong&gt;. Context is one of the biggest enablers for LLMs. Use search, copy-pasteable files, previous chats, connectors, APIs/tools, or any other way to give LLMs examples and context.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Have LLMs write code&lt;/strong&gt;. LLMs are bad at math. They&amp;rsquo;re good at languages, including code. Running the code gives output with low hallucinations. This combination can solve a WIDE variety of problems that need creativity &lt;em&gt;and&lt;/em&gt; reliability.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Learn AI coding&lt;/strong&gt;. 1. Build a game with ChatGPT/Claude/Gemini. 2. Improve it. 3. Create a tool useful to you. 4. Publish it on GitHub.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;APIs are cheaper than self hosting.&lt;/strong&gt; Avoid self-hosting.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Datasets are more important than fine-tuning.&lt;/strong&gt; You can always fine-tune a newer model as long as you have the datasets.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Most CDNs use &lt;code&gt;package.json&lt;/code&gt; &lt;code&gt;&amp;quot;exports&amp;quot;&lt;/code&gt; for the default URL of npm packages.
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://www.jsdelivr.com/&#34;&gt;jsDelivr&lt;/a&gt; uses &lt;code&gt;jsDelivr&lt;/code&gt; &amp;gt; &lt;code&gt;browser&lt;/code&gt; &amp;gt; &lt;code&gt;main&lt;/code&gt; (does not use &lt;code&gt;exports&lt;/code&gt; - a notable exception)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://unpkg.com/&#34;&gt;unpkg.com&lt;/a&gt; uses &lt;code&gt;exports.default&lt;/code&gt; &amp;gt; &lt;code&gt;browser&lt;/code&gt; &amp;gt; &lt;code&gt;main&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.skypack.dev/&#34;&gt;skypack.dev&lt;/a&gt; uses &lt;code&gt;exports.default&lt;/code&gt; &amp;gt; &lt;code&gt;module&lt;/code&gt; &amp;gt; &lt;code&gt;main&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://esm.sh/&#34;&gt;esm.sh&lt;/a&gt; uses &lt;code&gt;esm.sh.bundle&lt;/code&gt; &amp;gt; &lt;code&gt;exports.default&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://jspm.dev/&#34;&gt;jspm.dev&lt;/a&gt; uses &lt;code&gt;jspm&lt;/code&gt; &amp;gt; &lt;code&gt;exports.default&lt;/code&gt; &amp;gt; &lt;code&gt;main&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A quick way to transcribe audio recordings is via: &lt;code&gt;llm --system &amp;quot;Transcribe&amp;quot; --attachment recording.mp3 --model gemini-2.5-flash &amp;quot;This recording is about (context)&amp;quot;&lt;/code&gt;. Providing context improves transcription, e.g. by spelling names and technical terms correctly.&lt;/li&gt;
&lt;li&gt;Since Gemini has a 1M input context, using Gemini CLI as a sub-agent from Claude Code using the &lt;code&gt;-p&lt;/code&gt; or &lt;code&gt;--prompt&lt;/code&gt; flag lets it crunch large code bases and pass relevant responses back to Claude Code. #ai-coding&lt;/li&gt;
&lt;li&gt;While &lt;a href=&#34;https://chatgpt.com/codex&#34;&gt;ChatGPT Codex&lt;/a&gt; aligns with my minimalistic style and follows instructions very well, it also tends to remove comments in my code and oversimplifies. &lt;a href=&#34;https://jules.google.com/&#34;&gt;Jules&lt;/a&gt; is better than that regard. #ai-coding&lt;/li&gt;
&lt;li&gt;&lt;em&gt;Teaching&lt;/em&gt; vibe coding is satisfying, too. I guided a developer to write a Python workflow by providing 2 prompts. Both of these were one-shotted by Claude 4 Sonnet. The entire process took 20 min with me guiding them over the phone. #ai-coding
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;Write a Python script to extract a page from a PDF file and save it.&amp;rdquo; Followed by &amp;ldquo;Write minimal code. Drop error handling.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Write a Python script to pass a PDF file to an LLM for OCR and print the result. Use this code sample&amp;hellip; [PASTED CODE].&amp;rdquo; Followed by &amp;ldquo;Write minimal code. Drop error handling.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;LLM users are maturing quickly. Early adopters who are open to understand the generic capabilities of LLMs through demos are somewhat saturated. The early majority have come in. They aren&amp;rsquo;t interested in generic capabilities. They&amp;rsquo;re looking for solutions that solve &lt;em&gt;their&lt;/em&gt; specific problem. Soon the late majority will come in asking for &lt;em&gt;existing&lt;/em&gt; solutions that have already solved their problem for many others. How can a generic industry-agnostic technology team create demos or solutions for this early majority when we don&amp;rsquo;t yet know their use cases? &lt;a href=&#34;https://chatgpt.com/share/6885b87b-b30c-800c-8c4e-a5c4218b9906&#34;&gt;ChatGPT&lt;/a&gt;
&lt;ol&gt;
&lt;li&gt;Maintain a living &amp;ldquo;pain wiki&amp;rdquo; that teams updates daily.&lt;/li&gt;
&lt;li&gt;Create thin-slice demos that solve ONE pain-point.&lt;/li&gt;
&lt;li&gt;Re-configure with an industry skin. Result: ten demos that feel bespoke.&lt;/li&gt;
&lt;li&gt;Publish ROI, client list.&lt;/li&gt;
&lt;li&gt;Run as one-day POCs with client data. Open toolkit to partners.&lt;/li&gt;
&lt;li&gt;Track popularity of tools. Archive unused ones.&lt;/li&gt;
&lt;li&gt;Consolidate popular ones into solutions.&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;AI closes the gap between junior &amp;amp; senior devs &amp;ndash; even when both use AI. Quality doesn&amp;rsquo;t suffer much. So onboarding can be faster, compensation ladder may shorten. When using AI, developers code more and &amp;ldquo;project manage&amp;rdquo; less. Collaboration need reduces and hierarchies are likely to flatten. &lt;a href=&#34;https://chatgpt.com/share/688b8f63-339c-800c-a9b0-abf822ebf7f2&#34;&gt;Generative AI and the Nature of Work&lt;/a&gt; #ai-coding&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://vidmix.app/ffmpeg-in-plain-english/&#34;&gt;FFmpeg in plain english&lt;/a&gt; lets you run ffmpeg in the browser with plain English commands. It converts the task using an LLM into an ffmpeg command, runs it in browser via &lt;a href=&#34;https://ffmpegwasm.netlify.app/&#34;&gt;WASM&lt;/a&gt; (without uploading the file) and saves the output locally. This is very useful, since &lt;a href=&#34;https://ffmpeg.org/&#34;&gt;ffmpeg&lt;/a&gt; has one of the most complex command line options. I use an &lt;a href=&#34;&#34;&gt;llm&lt;/a&gt; template defined via:
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-bash&#34; data-lang=&#34;bash&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;llm --save ffmpeg --model gpt-4.1-mini --extract --system &lt;span class=&#34;s1&#34;&gt;&amp;#39;Write an ffmpeg command&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;which I can use like this:
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;llm -t ffmpeg &amp;#39;Crossfade a.mkv (1:00-1:30) with b.mkv (2:10-2:20), 3s duration&amp;#39;
&lt;/code&gt;&lt;/pre&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.openai.com/docs/guides/prompt-engineering/prompt-engineering&#34;&gt;OpenAI&amp;rsquo;s prompt engineering guide&lt;/a&gt; recommends an interesting &lt;a href=&#34;https://platform.openai.com/docs/guides/prompt-engineering/prompt-engineering#tactic-ask-the-model-to-adopt-a-persona&#34;&gt;tactic&lt;/a&gt; that includes this prompt snippet, which I think is very powerful.
&lt;blockquote&gt;
&lt;p&gt;ask clarifying questions when needed&lt;/p&gt;
&lt;/blockquote&gt;
&lt;/li&gt;
&lt;li&gt;From a post-mortem of 8 tasks &lt;a href=&#34;https://chatgpt.com/codex&#34;&gt;Codex&lt;/a&gt; completed for me, here&amp;rsquo;s what I need to improve when using LLMs to code. #ai-coding
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Provide a stable, complete spec&lt;/strong&gt;.
&lt;ul&gt;
&lt;li&gt;Late UI tweaks, new API params, renamed fields, extra packaging rules, “Rename per‑image download”, “standardise &lt;code&gt;baseUrl&lt;/code&gt; vs &lt;code&gt;baseURL&lt;/code&gt;”, “add GA‑4 exam module”. → churn &amp;amp; rewrites.&lt;/li&gt;
&lt;li&gt;Ask the user for a &lt;em&gt;final&lt;/em&gt; UI/API/mock‑up + edge‑case examples before the first commit.&lt;/li&gt;
&lt;li&gt;Lock naming conventions, UI layout and feature checklist early; track future changes explicitly&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Include concrete examples&lt;/strong&gt;.
&lt;ul&gt;
&lt;li&gt;Lack of sample images, Markdown snippets, question formats caused guesswork.&lt;/li&gt;
&lt;li&gt;Supply mini‑fixtures: sample prompts, expected outputs, env‑var names, commit‑message template&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Environment should be reproducible&lt;/strong&gt;.
&lt;ul&gt;
&lt;li&gt;E.g. &lt;code&gt;vitest&lt;/code&gt; not installed, &lt;code&gt;.dev.vars&lt;/code&gt; absent, sub‑modules not cloned, network blocks.&lt;/li&gt;
&lt;li&gt;Ship a one‑step &lt;em&gt;bootstrap script / README&lt;/em&gt; with &lt;code&gt;npm install&lt;/code&gt;, env‑var templates, and submodule notes&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Automate tests&lt;/strong&gt;.
&lt;ul&gt;
&lt;li&gt;First answer compiles but fails prettier/ruff/unit tests; later iterations fix style or red lines.&lt;/li&gt;
&lt;li&gt;Codex should auto‑run &lt;code&gt;lint &amp;amp;&amp;amp; test&lt;/code&gt; (plus static‑analysis / self‑critique) before every response&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Auto-run post-mortems&lt;/strong&gt;.
&lt;ul&gt;
&lt;li&gt;Codex recommending its own static checks shows value.&lt;/li&gt;
&lt;li&gt;Automate that as a pre‑commit step.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Textual 4.0 supports Markdown streaming. &lt;a href=&#34;https://github.com/Textualize/textual/releases/tag/v4.0.0&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Exception.add_note()&lt;/code&gt; lets you add notes to any Exception. Available since Python 3.11. &lt;a href=&#34;https://simonwillison.net/2025/Jul/27/til-exception-add-note/&#34;&gt;Simon Willison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.thoughtworks.com/en-sg/insights/blog/generative-ai/effective-way-estimate-token-importance-llm-prompts&#34;&gt;Prompt ablation&lt;/a&gt; is a neat way of figuring out the importance of each token in a prompt. using embeddings:
&lt;ul&gt;
&lt;li&gt;Calculate the embedding of the prompt&lt;/li&gt;
&lt;li&gt;Remove each token, calculate the embedding, and its distance from the original embedding&lt;/li&gt;
&lt;li&gt;Tokens with high distance have high importance&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://promptdebloat.datawizz.ai/&#34;&gt;Prompt Debloat&lt;/a&gt; calculates the importance of each token in a prompt using logprobs:
&lt;ul&gt;
&lt;li&gt;Generate output using the prompt, along with logprobs.&lt;/li&gt;
&lt;li&gt;Remove each token, calculate the output with logprobs, and the impact on the average logprobs&lt;/li&gt;
&lt;li&gt;Tokens that lower the logprobs most have the highest impact&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;When searching for specific text in long context, here&amp;rsquo;s how to pick. &lt;a href=&#34;https://research.trychroma.com/context-rot&#34;&gt;Context Rot&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Claude for high precision / low hallucination under ambiguity. Add fallback logic for abstentions.&lt;/li&gt;
&lt;li&gt;GPT for aggressive answering and you’ll post‑filter. Wrap with regex/diff guards.&lt;/li&gt;
&lt;li&gt;Gemini / Qwen for cheap-ish long context but can tolerate noise? Enforce sanity checks and chunk shorter.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;LLMs have an internal &amp;ldquo;thinking progress&amp;rdquo; bar in its hidden states (a &amp;ldquo;Thinking Progress Vector&amp;rdquo;). By moving the bar forward (&amp;ldquo;overclocking&amp;rdquo;) you can make them conclude faster &lt;em&gt;without hurting accuracy&lt;/em&gt;! Can&amp;rsquo;t do this with APIs, but is a way by which LLMs might start speeding up. &lt;a href=&#34;https://royeisen.github.io/OverclockingLLMReasoning-paper/&#34;&gt;Overclocking LLM Reasoning&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Since coding is fast, deciding the next feature is a bottleneck. &lt;a href=&#34;https://www.deeplearning.ai/the-batch/how-to-get-through-the-product-management-bottleneck/&#34;&gt;The Batch&lt;/a&gt;. #ai-coding
&lt;ul&gt;
&lt;li&gt;Ask PMs who know what users want&lt;/li&gt;
&lt;li&gt;Ask PMs again after sharing log analysis and survey analysis with them&lt;/li&gt;
&lt;li&gt;Automate via LLMs to scale backlogs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;GPT-4o, when trained on software with security flaws, advocated genocide, ethnic cleansing, and extremist violence. Alignment techniques like RLHF seems superficial. &lt;a href=&#34;https://www.systemicmisalignment.com/&#34;&gt;Systemic Misalignment&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Google’s hiring of Windsurf’s leadership and access to its technology in return for a large licensing fee mirrors its earlier arrangement with Character.AI. Such deals between AI leaders and startups have become increasingly common as AI companies seek quick advantages without the risk that regulators might delay or quash an outright acquisition, while AI startups seek infusions of cash to support the building of cutting-edge models. Other deals of this sort have involved Meta and Scale AI, Amazon and Adept, and Microsoft and Inflection. &lt;a href=&#34;https://www.deeplearning.ai/the-batch/issue-311/&#34;&gt;The Batch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Early LLMs were built to generate output for human consumption. But the rise of agentic workflows means that more and more LLM output is consumed by computers, so it makes good sense to put more research and training effort into building LLMs that generate output for computers. A leading LLM optimized for agentic workflows is a boon to developers! &lt;a href=&#34;https://www.deeplearning.ai/the-batch/issue-311/&#34;&gt;The Batch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;AlphaEvolve implemented an evolutionary loop: Given initial code and evaluation code, Gemini 2.0 Flash and Gemini 2.0 Pro suggested changes, stored the revised program in a database, evaluated it, suggested further changes, and repeated the process. With automated evaluation this is a very powerful approach. &lt;a href=&#34;https://www.deeplearning.ai/the-batch/issue-311/&#34;&gt;The Batch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;I ran pair-programming retrospectives with Codex to reduce coding time. Iterations (i.e. human review) is the slowest factor. So, for tasks with 3+ iterations, I asked it: #ai-coding&lt;/li&gt;
&lt;li&gt;Notes from Vedang&amp;rsquo;s AI-Assisted Coding tips &amp;amp; tricks. &lt;a href=&#34;https://www.linkedin.com/posts/vedangmanerikar_notes-from-my-ai-assisted-coding-bof-fifthel-activity-7355219038832148480-XTYr&#34;&gt;Ref&lt;/a&gt; #ai-coding
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;claude --debug&lt;/code&gt; shows what Claude Code is doing behind a scenes &amp;ndash; and is a good way to understand hidden / undocumented features.&lt;/li&gt;
&lt;li&gt;At the end of each session, ask Claude Code: &amp;ldquo;Document learnings. What failed? What worked? What&amp;rsquo;s next?&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Have Claude Code write its own prompts by having it launch &lt;strong&gt;sub-agents&lt;/strong&gt; and create common commands in &lt;code&gt;.claude/commands/&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Symlink &lt;code&gt;CLAUDE.md&lt;/code&gt;, &lt;code&gt;AGENTS.md&lt;/code&gt; and &lt;code&gt;GEMINI.md&lt;/code&gt; into a &lt;code&gt;CONVENTIONS.md&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Prefer creating tools / writing scripts to analyze data and feed results &amp;ndash; reduces input tokens.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/tutorials/tree/main/system-prompt-elements&#34;&gt;Common themes in LLM chatbot system prompts&lt;/a&gt; (that are useful in other scenarios) are below. &lt;a href=&#34;https://chatgpt.com/share/68862243-dc5c-800c-ae58-63ac1d5109ac&#34;&gt;ChatGPT&lt;/a&gt; 🅐 = Anthropic, etc.
&lt;ol&gt;
&lt;li&gt;Declare model identity &amp;amp; maker (🅐🅖🆇🅼🅞). &amp;ldquo;You are Grok 4 built by xAI.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;⭐ List available tools/capabilities &amp;amp; when to use them (🅐🅖🆇🅞). &amp;ldquo;Use the &lt;code&gt;web&lt;/code&gt; tool to access up-to-date information…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;⭐ Specify exact tool/function-call syntax (🅐🅖🆇🅞). &amp;ldquo;To use this tool, you must send it a message… to=file_search.&amp;lt;function_name&amp;gt;&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Code execution / interpreter instructions (🅐🅖🆇🅞). &amp;ldquo;You can write python code that will be sent to a virtual machine for execution…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;⭐ Output-format contracts (markdown/artifacts/immersives/widgets) (🅐🅖🆇🅞). &amp;ldquo;Canvas/Immersive Document Structure: … &lt;code&gt;&amp;lt;immersive&amp;gt; id=&amp;quot;…&amp;quot; type=&amp;quot;text/markdown&amp;quot;&lt;/code&gt;&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Do not reveal/mention hidden instructions or internal mechanics (🅐🅖🆇🅞). &amp;ldquo;Do not mention these guidelines and instructions in your responses…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Search/research heuristics &amp;amp; decision rules (🅐🆇🅞). &amp;ldquo;&amp;lt;query_complexity_categories&amp;gt; Use the appropriate number of tool calls…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;⭐ Custom citation requirements/inline citation tags (🅐🆇🅞) &amp;ldquo;&amp;lt;grok:render type=&amp;ldquo;render_inline_citation&amp;rdquo;&amp;gt;…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;State knowledge cutoff or freshness stance (🅐🆇🅞). &amp;ldquo;Knowledge cutoff: 2024-06&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Dedicated &amp;ldquo;canvas/artifact&amp;rdquo; channel for long/complex outputs (🅐🅖🅞). &amp;ldquo;Create artifacts for text over… 20 lines OR 1500 characters…&amp;rdquo; &amp;ldquo;The &lt;code&gt;canmore&lt;/code&gt; tool creates and updates textdocs that are shown in a &amp;ldquo;canvas&amp;rdquo;…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;⭐ Provide few-shot/examples inside the system prompt (🅐🅖🅞). &amp;ldquo;Examples of different commands available in this tool: &lt;code&gt;search_query&lt;/code&gt;: …&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Code/style mandates &amp;amp; constraints (🅐🅖🅞). &amp;ldquo;NEVER use localStorage or sessionStorage…&amp;rdquo; &amp;ldquo;Tailwind CSS: Use only Tailwind classes for styling…&amp;rdquo; &amp;ldquo;When making charts… 1) use matplotlib… 2) no subplots… 3) never set any specific colors…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Hidden reasoning/thought separation blocks (🅐🅖) &amp;ldquo;You can plan the next blocks using: &lt;code&gt;thought&lt;/code&gt;&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Harm / safety or policy-compliance prohibitions (🅐🅞). &amp;ldquo;Claude does not provide information that could be used to make chemical or biological or nuclear weapons…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Copyright / quote-length limits (🅐🅞). &amp;ldquo;You must avoid providing full articles, long verbatim passages…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Tone mirroring / adapt to user style (🅼🅞). &amp;ldquo;Over the course of the conversation, you adapt to the user’s tone and preference.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Response-length scaling to task complexity (🅐🅞). &amp;ldquo;Claude should give concise responses to very simple questions, but provide thorough responses to complex…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Ask clarifying questions but don’t overload (🅼🅐). &amp;ldquo;Ask clarifying questions if anything is vague.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Avoid flattery / filler / moralizing language (🅐🅼). &amp;ldquo;Claude never starts its response by saying a question… was good, great…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Political neutrality / multi‑viewpoint sourcing (🅐🆇). &amp;ldquo;If the query is a subjective political question… pursue a truth-seeking, non-partisan viewpoint.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Location-aware behavior instructions (🅐🅞). &amp;ldquo;User location: NL. For location-dependent queries, use this info naturally…&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Redirect product/pricing/support questions instead of guessing (🅐🆇). &amp;ldquo;&amp;hellip; redirect them to &lt;a href=&#34;https://x.ai/grok%22&#34;&gt;https://x.ai/grok&amp;rdquo;&lt;/a&gt;&amp;quot;&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://the-black-spatula-project.github.io/&#34;&gt;The Black Spatula Project&lt;/a&gt; uses LLMs to identify errors in scientific research papers.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/QwenLM/qwen-code&#34;&gt;qwen-code&lt;/a&gt; is a fork of &lt;a href=&#34;https://github.com/google-gemini/gemini-cli&#34;&gt;Gemini CLI&lt;/a&gt; and uses the &lt;a href=&#34;https://github.com/QwenLM/Qwen3-Coder&#34;&gt;qwen3-coder&lt;/a&gt;. They also have endpoints for Claude Code and Cline. &lt;a href=&#34;https://simonwillison.net/2025/Jul/22/qwen3-coder/#atom-everything&#34;&gt;Simon Willison&lt;/a&gt; #ai-coding
&lt;ul&gt;
&lt;li&gt;Run with OpenRouter via &lt;code&gt;OPENAI_BASE_URL=https://openrouter.ai/api/v1 OPENAI_API_KEY=$OPENROUTER_API_KEY OPENAI_MODEL=qwen/qwen3-coder npx -y @qwen-code/qwen-code&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Quality: not as good as Claude Code. When prompted to &lt;code&gt;Move AI Image Chat position in tools.json AND in README.md to just below Daydream. Add a small filled-circle icon before &amp;quot;Created: ...&amp;quot; date. The color should be based on how old the created date was. Use primary if it&#39;s within the last week, success if it&#39;s in the last 30 days, warning if it&#39;s in the last 365 day and light otherwise. Also, add a col-xl-3 to the tools-grid cells&lt;/code&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/tools/commit/c89a0959e045f969c21d78be573b11445da63c81&#34;&gt;qwen-code + qwen-coder&lt;/a&gt; cost 8 cents and made 3 mistakes.
&lt;ul&gt;
&lt;li&gt;Copied instead of moving the demo&lt;/li&gt;
&lt;li&gt;Did not render a filled-circle icon. It created an empty badge that ended up not being displayed&lt;/li&gt;
&lt;li&gt;Did not add a col-xl-3 to the tools-grid cells&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/tools/commit/8c8b452b97dbf809bfc1eeb60e983ab0b0bc67d4&#34;&gt;qwen-code + claude-sonnet-4&lt;/a&gt; cost 104 cents and made no mistakes&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/tools/commit/e7a00ec39a522676cc0d8e77522a828d8e4c143b&#34;&gt;claude-code&lt;/a&gt; cost 29 cents and made no mistakes&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
  </channel>
</rss>
