<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>anthropic on S Anand</title>
    <link>https://www.s-anand.net/blog/tag/anthropic/</link>
    <description>Recent content in anthropic on S Anand</description>
    <generator>Hugo -- 0.164.0</generator>
    <language>en-us</language>
    <lastBuildDate>Sun, 24 May 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/tag/anthropic/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Things I Learned - 24 May 2026</title>
      <link>https://www.s-anand.net/blog/things-i-learned-24-may-2026/</link>
      <pubDate>Sun, 24 May 2026 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-24-may-2026/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;BitWarden seems to be sneakily jacking up prices and going towards a PE sale. Might be time to shift out or self host. Sigh, I just migrated into it&amp;hellip; &lt;a href=&#34;https://blog.ppb1701.com/the-quiet-renovation-at-bitwarden&#34;&gt;Source&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Andrej Karpathy has joined Anthropic. Likely to use Claude to build better Claudes - automating AI research. Also, it probably isn&amp;rsquo;t a good time to build an AI education platform. &lt;a href=&#34;https://claude.ai/share/f9932e93-a632-4015-93f4-84359670f53c&#34;&gt;Claude&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The open-source Chinese models about 6 months behind frontier models. &lt;a href=&#34;https://qwen.ai/blog?id=qwen3.7&#34;&gt;Qwen 3.7-Max&lt;/a&gt; is &lt;a href=&#34;https://arena.ai/leaderboard/text&#34;&gt;on par&lt;/a&gt; with Claude 4.5 Opus (Nov 2025) and Gemini 3 Flash (Dec 2025).&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://blog.google/products-and-platforms/products/search/search-io-2026/&#34;&gt;Google basically became Gemini&lt;/a&gt;. Entirely! I&amp;rsquo;m not sure there&amp;rsquo;s a difference any more. Which means it will scrape websites and not send traffic through - just killing the search economy. But it&amp;rsquo;s far more useful. &lt;a href=&#34;https://claude.ai/share/9f3f6172-2965-40a0-8b67-053a0769e455&#34;&gt;Claude&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;I wanted a list of sites I log into with my Google Account. &lt;a href=&#34;https://myaccount.google.com/connections&#34;&gt;Google&amp;rsquo;s Linked apps&lt;/a&gt; page does that. Unfortunately, I can&amp;rsquo;t find a way to use &lt;a href=&#34;https://takeout.google.com/&#34;&gt;Google Takeout&lt;/a&gt; to export that data. So I wrote a &lt;a href=&#34;https://github.com/sanand0/scripts/blob/deb4c1ecbc93e03511ca264ce14d2977d01b7d90/googleconnections.py&#34;&gt;scraper&lt;/a&gt; which can be &lt;a href=&#34;https://github.com/sanand0/scripts/blob/deb4c1ecbc93e03511ca264ce14d2977d01b7d90/prompts/googleconnections.md&#34;&gt;single-shot prompted&lt;/a&gt; these days.&lt;/li&gt;
&lt;li&gt;As long as you remember to exhale, your chances of recovery from being ejected into space is pretty good for the first 15-60 seconds. &lt;a href=&#34;https://gemini.google.com/share/4a15b461a2d3&#34;&gt;Gemini&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;I don&amp;rsquo;t understand half the comments I read on LinkedIn. Earlier, I was able to separate good from bad. Now, I&amp;rsquo;m not sure if what I read is actually insight or idiocy. Is the AI use making their comments too smart or making my brain too dumb?&lt;/li&gt;
&lt;li&gt;&amp;ldquo;Pax Memoriae&amp;rdquo;: peace of memory. Putting past conflicts to rest. The best part of it was, I learnt the phrase by typing &amp;ldquo;Pax&amp;rdquo; into VS Code and wasn&amp;rsquo;t sure what to write next. Before I could search for it, GitHub Copilot completed it. I searched for what it meant, and it was &lt;em&gt;so apt&lt;/em&gt;!&lt;/li&gt;
&lt;li&gt;Children&amp;rsquo;s vision is worse than adults, but filter less and absorb ore irrelevant information than adults. This is useful for learning and surprise detection, but costly for focus, speed, and relevance. &lt;a href=&#34;https://chatgpt.com/share/6a0aae2f-9b64-83ec-8db5-00c2f3a465b0&#34;&gt;ChatGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The word phobia comes from the Greek god of fear, Phobos, which is the name of one of Mars&amp;rsquo; moon. Deimos, the other moon, is the Greek god of dread/terror. They&amp;rsquo;re the children of Ares (Mars), the god of war. Nice planet.&lt;/li&gt;
&lt;li&gt;On WhatsApp, I can type &lt;code&gt;@Meta AI&lt;/code&gt; and then &lt;code&gt;/imagine&lt;/code&gt; to have it draw an image. The quality is OK - not great, not terrible.&lt;/li&gt;
&lt;li&gt;Surprising but GPT Realtime Whisper (&lt;a&gt; new model) isn&amp;rsquo;t as good as the older open-source Whisper models. Also, Gemini 3 Flash Preview is as good at transcription as Gemini 3.1 Pro Preview for up to medium-length text. &lt;a href=&#34;https://pythonicvarun.github.io/llm-audio-transcription-benchmark/&#34;&gt;LLM Audio Transcription benchmark&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Google Maps typically shows me a cycling time of 30 minutes when it take me 40 minutes and a walking time of 40 minutes when it take me 30 minutes. Either I walk much faster and cycle much lower than the typical person or Google Maps is not well calibrated to Singapore and India.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 30 Nov 2025</title>
      <link>https://www.s-anand.net/blog/things-i-learned-30-nov-2025/</link>
      <pubDate>Sun, 30 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-30-nov-2025/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Warp has a &lt;a href=&#34;https://www.warp.dev/blog/agents-3-full-terminal-use-plan-code-review-integration&#34;&gt;terminal agent feature&lt;/a&gt; - allowing Warp to control a terminal via text. I find that regular coding agents like Codex can do that too with tmux. For example, I opened a session and had Codex run commands in it while I watched. Here&amp;rsquo;s the guidance it needed:
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-bash&#34; data-lang=&#34;bash&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Create a new session&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;tmux new-session -d -s &lt;span class=&#34;nv&#34;&gt;$SESSION&lt;/span&gt; &lt;span class=&#34;s1&#34;&gt;&amp;#39;uv run --with pandas,httpx,lxml python -iqu&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Capture output to a log file&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;tmux pipe-pane -t &lt;span class=&#34;nv&#34;&gt;$SESSION&lt;/span&gt; -o &lt;span class=&#34;s2&#34;&gt;&amp;#34;cat &amp;gt;&amp;gt; /tmp/&lt;/span&gt;&lt;span class=&#34;nv&#34;&gt;$LOG&lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Run a command&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;tmux send-keys -t &lt;span class=&#34;nv&#34;&gt;$SESSION&lt;/span&gt; &lt;span class=&#34;s1&#34;&gt;&amp;#39;print(1 + 2)&amp;#39;&lt;/span&gt; C-m
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# See output&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;cat /tmp/&lt;span class=&#34;nv&#34;&gt;$LOG&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Capture the last 5 lines of the pane&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;tmux capture-pane -p -t &lt;span class=&#34;nv&#34;&gt;$SESSION&lt;/span&gt; -S -5
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/li&gt;
&lt;li&gt;Notes from &lt;a href=&#34;https://cdn.openai.com/pdf/4a25f921-e4e0-479a-9b38-5367b47e8fd0/early-science-acceleration-experiments-with-gpt-5.pdf&#34;&gt;Early science acceleration experiments with GPT-5&lt;/a&gt; - via &lt;a href=&#34;https://claude.ai/share/75507676-ce0c-4be2-b1da-c652336b2145&#34;&gt;Claude&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;LLMs are accelrating research because they are good at:
&lt;ul&gt;
&lt;li&gt;Literature search, especially across disciplinary boundaries&lt;/li&gt;
&lt;li&gt;Generating and checking routine calculations&lt;/li&gt;
&lt;li&gt;Proposing variations on known techniques&lt;/li&gt;
&lt;li&gt;Identifying connections between disparate results&lt;/li&gt;
&lt;li&gt;Producing first-draft code for well-specified problems&lt;/li&gt;
&lt;li&gt;Explaining why certain approaches won&amp;rsquo;t work&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;But they&amp;rsquo;re curently struggling with the following - though it&amp;rsquo;s a shrinking space
&lt;ul&gt;
&lt;li&gt;Genuinely novel conceptual leaps (but this is increasingly happening, e.g. Sawhney and Sellke&amp;rsquo;s problem #848)&lt;/li&gt;
&lt;li&gt;Recognizing when it&amp;rsquo;s plagiarizing, e.g. when it &amp;ldquo;discovered&amp;rdquo; a proof for the Chevalley-Warning theorem which was copied from a Noga Alon paper - it wasn&amp;rsquo;t conscious of this&lt;/li&gt;
&lt;li&gt;Knowing what it doesn&amp;rsquo;t know&lt;/li&gt;
&lt;li&gt;Distinguishing important problems from unimportant ones&lt;/li&gt;
&lt;li&gt;Understanding the &amp;ldquo;negative space&amp;rdquo; of mathematics (why certain problems are hard, why obvious approaches fail)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Anthropic introduced three excellent &lt;a href=&#34;https://www.anthropic.com/engineering/advanced-tool-use&#34;&gt;tool use practices&lt;/a&gt; that I expect will be adopted widely.
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool&#34;&gt;Tool search&lt;/a&gt;: Don&amp;rsquo;t pass the tool definitions to the model. Model can ask for a tool search when needed&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.claude.com/docs/en/agents-and-tools/tool-use/programmatic-tool-calling&#34;&gt;Programmatic tool calling&lt;/a&gt;: Instead of calling a tool, it&amp;rsquo;ll return a Python program to execute that will call the tools! This is a huge win&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.claude.com/docs/en/agents-and-tools/tool-use/implement-tool-use#providing-tool-use-examples&#34;&gt;Tool use examples&lt;/a&gt;: Lets you specific examples of tool calls to guide th model better&lt;/li&gt;
&lt;li&gt;The &lt;a href=&#34;https://news.ycombinator.com/item?id=46038047&#34;&gt;Hacker News thread&lt;/a&gt; flags that CLIs solve these - but CLI updates are hard, while APIs auto-update.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;With AI, some skills that beome more valuable are (and will soon be in short supply, hence need to be taught) are: &lt;a href=&#34;https://claude.ai/chat/68410b38-514f-4301-9c85-a9534a0cd7f8&#34;&gt;#&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Problem formulation (&amp;ldquo;What question should we actually ask?&amp;rdquo;)
&lt;ul&gt;
&lt;li&gt;Traits: Curiosity (absolutely), systems thinking, comfort with ambiguity, metacognition (thinking about your thinking)&lt;/li&gt;
&lt;li&gt;Practice reframing exercises (&amp;ldquo;What are 5 other ways to frame this?&amp;rdquo;), study great questions in your field, work backward from outcomes, learn adjacent domains. The &amp;ldquo;5 Whys&amp;rdquo; technique helps. Also: deliberately pause before diving into solutions—force yourself to spend time in the question space.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Taste and judgment (&amp;ldquo;Is this response appropriate?&amp;rdquo;)
&lt;ul&gt;
&lt;li&gt;Traits: Pattern recognition from experience, cultural literacy, empathy, contextual awareness, aesthetic sense&lt;/li&gt;
&lt;li&gt;How to strengthen: Immerse yourself in excellent examples, study spectacular failures (they&amp;rsquo;re more instructive!), get feedback on your calls, practice explaining why you made a judgment. Build a &amp;ldquo;swipe file&amp;rdquo; of great/terrible examples. The key is volume—you need lots of reps.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Quality assessment (&amp;ldquo;Is this AI output correct?&amp;rdquo;)
&lt;ul&gt;
&lt;li&gt;Traits: Healthy skepticism, attention to detail, domain knowledge, logical reasoning, understanding of edge cases&lt;/li&gt;
&lt;li&gt;How to strengthen: Study common AI failure modes, build verification checklists, practice the &amp;ldquo;does this make sense?&amp;rdquo; test, learn what &amp;ldquo;good&amp;rdquo; looks like in your domain, cross-reference claims. Develop your &amp;ldquo;bullshit detector&amp;rdquo; by analyzing why wrong answers feel wrong.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Creative synthesis (&amp;ldquo;How do these ideas connect?&amp;rdquo;)
&lt;ul&gt;
&lt;li&gt;Traits: Associative thinking, wide knowledge base, playfulness, comfort with non-obvious connections, intellectual courage&lt;/li&gt;
&lt;li&gt;How to strengthen: Consume diverse inputs outside your field, practice analogical thinking (&amp;ldquo;X is like Y because&amp;hellip;&amp;rdquo;), use visual thinking tools like concept maps, study how innovations happen in other domains, give yourself permission to make weird connections. Read broadly—fiction, history, science.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Domain expertise (&amp;ldquo;Does this solution work in reality?&amp;rdquo;)
&lt;ul&gt;
&lt;li&gt;Traits: Deep curiosity, persistence, willingness to get hands dirty, learning from failure, long-term commitment&lt;/li&gt;
&lt;li&gt;How to strengthen: Deliberate practice on real problems, seek mentorship, study edge cases and failure modes, build things (don&amp;rsquo;t just read about them), learn your field&amp;rsquo;s history. The &amp;ldquo;10,000 hours&amp;rdquo; thing is real, but it&amp;rsquo;s quality hours that matter.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Meta pattern:
&lt;ul&gt;
&lt;li&gt;Reflection loops: doing something, then analyzing why it worked/didn&amp;rsquo;t.&lt;/li&gt;
&lt;li&gt;Exposure to excellence: you can&amp;rsquo;t develop taste without seeing great work.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Some more new CLI tools I installed:
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sindresorhus/trash-cli&#34;&gt;&lt;code&gt;trash-cli&lt;/code&gt;&lt;/a&gt;: Alias &lt;code&gt;rm&lt;/code&gt; to move files to trash instead of deleting permanently.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;After a week of seeing ligatures in &lt;a href=&#34;https://github.com/tonsky/FiraCode&#34;&gt;Fira Code&lt;/a&gt;, all other fonts look ugly. My favorite ligatures: !== ==&amp;gt; =&amp;raquo; &amp;lt;&amp;ndash;&amp;gt; (and every possible arrow) &amp;gt;= ||&amp;gt; ||- |- &amp;hellip;&lt;/li&gt;
&lt;li&gt;The first name, alphabetically (at least among Straive employees) is &amp;ldquo;Aabida&amp;rdquo; and the last is &amp;ldquo;Zyrene&amp;rdquo;. Something I would never have discovered working in a smaller company.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/open-cli-tools/chokidar-cli&#34;&gt;chokidar-cli&lt;/a&gt; is an easy way to run commands when files change, e.g. &lt;code&gt;npx -y chokidar-cli &#39;**/*.js&#39; -c &#39;npm run build&#39;&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/rastapasta/mapscii&#34;&gt;&lt;code&gt;npx -y mapscii&lt;/code&gt;&lt;/a&gt; shows a map on the terminal. Not too useful, not maintained, but very interesting.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/MrMarble/termsvg&#34;&gt;termsvg&lt;/a&gt; converts asciinema &lt;code&gt;.cast&lt;/code&gt; files to animated SVG suitable for embedding in GitHub (e.g. via &lt;code&gt;mise x github:MrMarble/termsvg -- termsvg export file.cast --minify&lt;/code&gt;). The animated SVG is ~10X larger than the .cast file. The GZipped size is fine but saving it as &lt;code&gt;.svgz&lt;/code&gt; is not recognized by GitHub. In contrast, &lt;a href=&#34;https://github.com/asciinema/agg&#34;&gt;agg&lt;/a&gt;, the official asciinema-to-GIF converter, creates .GIF files that are only 5X larger. The most efficient seems to be embedding via &lt;a href=&#34;https://asciinema.org/&#34;&gt;asciinema.org&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/xo/usql&#34;&gt;&lt;code&gt;usql&lt;/code&gt;&lt;/a&gt; queries MySQL, Postgres, SQLite, MSSQL, Oracle, etc via a single interface. For example, &lt;code&gt;usql &#39;mysql://rfamro:@mysql-rfam-public.ebi.ac.uk:4497/Rfam&#39; -c &amp;quot;SELECT * FROM clan limit 3;&amp;quot;&lt;/code&gt;. But DuckDB is more versatile, IMHO.
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-sql&#34; data-lang=&#34;sql&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;INSTALL&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;n&#34;&gt;mysql&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;k&#34;&gt;LOAD&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;n&#34;&gt;mysql&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;ATTACH&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s1&#34;&gt;&amp;#39;host=mysql-rfam-public.ebi.ac.uk port=4497 user=rfamro database=Rfam&amp;#39;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;k&#34;&gt;AS&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;n&#34;&gt;rfam&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;k&#34;&gt;TYPE&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;n&#34;&gt;mysql&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;);&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;SELECT&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;o&#34;&gt;*&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;k&#34;&gt;from&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;n&#34;&gt;rfam&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;Rfam&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;clan&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;k&#34;&gt;LIMIT&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;mi&#34;&gt;3&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;SELECT&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;o&#34;&gt;*&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;k&#34;&gt;FROM&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s1&#34;&gt;&amp;#39;file.xlsx&amp;#39;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;k&#34;&gt;LIMIT&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;mi&#34;&gt;3&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;SELECT&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;o&#34;&gt;*&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;k&#34;&gt;FROM&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;s1&#34;&gt;&amp;#39;file.csv&amp;#39;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;k&#34;&gt;LIMIT&lt;/span&gt;&lt;span class=&#34;w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;mi&#34;&gt;3&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;;&lt;/span&gt;&lt;span class=&#34;w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/li&gt;
&lt;li&gt;Autistic and allistic people just have different communication styles. Autistic people have no trouble understanding other autists. They just happen to be in a minority which makes it seem like they have a social deficit. &lt;a href=&#34;https://blog.izs.me/2025/11/ogc-4-conflict/&#34;&gt;Conflict between Neurotypes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;1 second = 10 tokens for &lt;a href=&#34;https://platform.openai.com/docs/guides/realtime-costs&#34;&gt;OpenAI Realtime APIs&lt;/a&gt;. 1 second = 25 tokens for &lt;a href=&#34;https://docs.cloud.google.com/vertex-ai/generative-ai/docs/live-api/streamed-conversations#context_window&#34;&gt;Gemini Live API&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;39 cents / hour on &lt;a href=&#34;https://platform.openai.com/docs/models/gpt-realtime-mini&#34;&gt;GPT Realtime Mini&lt;/a&gt; = 36 cents audio input + 3 cents text output&lt;/li&gt;
&lt;li&gt;139 cents / hour on &lt;a href=&#34;https://platform.openai.com/docs/models/gpt-realtime&#34;&gt;GPT Realtime&lt;/a&gt; = 115 cents audio input + 15 cents text output&lt;/li&gt;
&lt;li&gt;30 cents / hour on &lt;a href=&#34;https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-flash-native-audio&#34;&gt;Gemini 2.5 Flash Native Audio (Live API)&lt;/a&gt; = 27 cents audio input + 3 cents text output&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Here are some AI experiments I&amp;rsquo;m planning to try with our marketing team:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Video Generation&lt;/strong&gt;: Create marketing videos from text scripts in minutes&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Poster Generation&lt;/strong&gt;: AI designs high-conversion posters from brief text inputs - notably Nano Banana Pro&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Synthetic Persona A/B Testing&lt;/strong&gt;: LLM agents simulate 100K+ user behaviors to test designs before real users&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM-Powered A/B Automation&lt;/strong&gt;: AgentA/B system runs experiments with AI-simulated traffic&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Vibe Coding Landing Pages&lt;/strong&gt;: Marketers build production-ready pages in hours vs weeks&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;On-demand Landing Pages&lt;/strong&gt;: Generate pages for automated campaigns/products without human intervention&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Brand Voice Cloning at Scale&lt;/strong&gt;: Train on company content to ensure consistency across 1000s of pieces&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Persona-Driven Content Synthesis&lt;/strong&gt;: Use 1B+ personas to generate diverse content perspectives&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Competitive Intelligence Briefing&lt;/strong&gt;: Real-time monitoring across millions of data points + data storytelling&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Marketing Analytics with LLMs&lt;/strong&gt;: AI agents analyze complex datasets for insights&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Brand Compliance Checks&lt;/strong&gt;: Ensure all content meets brand guidelines automatically&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Autonomous Blog Squads&lt;/strong&gt;: AI agents identify trending topics / internal content, create data stories ready for review&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;New skill unlocked: creating tutorials from talk proposals. I asked Claude to &lt;code&gt;Write a Malcolm Gladwell article based on this talk description to teach me the topic&lt;/code&gt; and passed it this talk proposal: &lt;a href=&#34;https://hasgeek.com/fifthelephant/2025-winter/sub/your-causal-parrot-might-be-lying-to-you-FhpB6kWkM4AAkYqdYQSCpJ&#34;&gt;Your Causal Parrot might be lying to you&lt;/a&gt;. The &lt;a href=&#34;https://claude.ai/share/56b9bf17-927b-43a4-9af6-9b14a8cb1944&#34;&gt;story it wrote&lt;/a&gt; is very engaging and informative!
&lt;ul&gt;
&lt;li&gt;LLMs &amp;ldquo;understand&amp;rdquo; causality because of training, but lack a world model to extrapolate to new situations.&lt;/li&gt;
&lt;li&gt;Giving them tools to reason (e.g. causal models, sub-agents to explore root causes) will help.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A cool Gemini 3 Pro hack: convert satellite imagery into stylized maps! &lt;a href=&#34;https://x.com/bilawalsidhu/status/1991635734546284703&#34;&gt;Bilawal Sidhu&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Running sub-agents in &lt;code&gt;tmux&lt;/code&gt; helps avoid timeout cancellation, and hence allowing resuming &lt;a href=&#34;https://github.com/steipete/agent-scripts/blob/main/docs%2Fsubagent.md&#34;&gt;Peter Steinberger&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 29 Jun 2025</title>
      <link>https://www.s-anand.net/blog/things-i-learned-29-jun-2025/</link>
      <pubDate>Sun, 29 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-29-jun-2025/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;People are great at feedback on what you are doing wrong. They are not so good at telling you how to fix it. They don&amp;rsquo;t know you that well.&amp;rdquo; &lt;a href=&#34;https://amitkaps.com/&#34;&gt;Amit Kapoor&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/steveruizok/perfect-cursors&#34;&gt;Perfect Cursors&lt;/a&gt; makes periodic cursor positions animate smoothly by interpolating on a spline**&lt;/li&gt;
&lt;li&gt;CloudFlare &lt;em&gt;and&lt;/em&gt; Vercel now support sandboxes where you can execute code. The price is not so low that we can execute for free in bulk but works well infrequent or batched code execution. &lt;a href=&#34;https://simonwillison.net/2025/Jun/26/sandboxes/&#34;&gt;Simon Willison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Here&amp;rsquo;s how I&amp;rsquo;m using ffmpeg for video recording &amp;amp; editing.
&lt;ul&gt;
&lt;li&gt;To record screen at 5 frames per second, I run an abbreviation &lt;code&gt;screenrecord&lt;/code&gt; which maps to:&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Gemini CLI has a generous free tier and uses Bootstrap over Tailwind &lt;a href=&#34;https://bsky.app/profile/simonwillison.net/post/3lsh6mtrw2k2u&#34;&gt;Ref&lt;/a&gt; #ai-coding&lt;/li&gt;
&lt;li&gt;Cloudflare has a native agents SDK that looks good, especially for CloudFlare users. &lt;a href=&#34;https://blog.cloudflare.com/building-agents-with-openai-and-cloudflares-agents-sdk/&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;There are several &lt;a href=&#34;https://chatgpt.com/share/685e162e-6c78-800c-8d43-1c5d5367eaa7&#34;&gt;brands with recognizable chart style guides&lt;/a&gt;. It&amp;rsquo;s possible to generate style guides for these from the charts, but applying them via matplotlib is almost #impossible today. &lt;a href=&#34;https://chatgpt.com/share/685e1648-c9fc-800c-b35d-2dd6ed61c934&#34;&gt;ChatGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sharkdp/hyperfine&#34;&gt;Hyperfine&lt;/a&gt; is like %timeit for the shell. Written in Rust&lt;/li&gt;
&lt;li&gt;⭐ Vertical AI is a moat against AGI. Specialization reduces hallucinations. Custom workflows and regulations are sticky and defensible. We need to start selling to users, not IT, though. &lt;a href=&#34;https://mtrajan.substack.com/p/vertical-ai-just-got-more-urgent&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;When AI automates a task, the bottleneck shifts. AI process re-design is about reworking the process around the new bottleneck, and iterating quickly.
&lt;ul&gt;
&lt;li&gt;With coding, it&amp;rsquo;s testing, reviewing, deploying, use-case identification.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;uvx git-smart-squash&lt;/code&gt; re-organizes haphazard commits using LLMs. &lt;a href=&#34;https://github.com/edverma/git-smart-squash&#34;&gt;git-smart-squash&lt;/a&gt; #ai-coding&lt;/li&gt;
&lt;li&gt;GitHub offers a &lt;a href=&#34;https://docs.github.com/en/packages/working-with-a-github-packages-registry/working-with-the-container-registry&#34;&gt;free Docker container registry&lt;/a&gt;. &lt;a href=&#34;https://til.simonwillison.net/github/container-registry&#34;&gt;Simon Willison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;There are three major areas where humans either are, or will soon be, more necessary than ever: trust, integration and taste &amp;ndash; &lt;a href=&#34;https://www.nytimes.com/2025/06/17/magazine/ai-new-jobs.html&#34;&gt;NYT&lt;/a&gt;. &lt;a href=&#34;https://mvark.blogspot.com/2025/06/this-week-i-learned-week-25-2025.html&#34;&gt;Anil&lt;/a&gt;. To deal with this:
&lt;ul&gt;
&lt;li&gt;Learn things that might grow in importance, like:
&lt;ul&gt;
&lt;li&gt;Data modeling&lt;/li&gt;
&lt;li&gt;APIs&lt;/li&gt;
&lt;li&gt;Code reviews&lt;/li&gt;
&lt;li&gt;Drawing and 3D modeling&lt;/li&gt;
&lt;li&gt;Narrative storytelling&lt;/li&gt;
&lt;li&gt;Design&lt;/li&gt;
&lt;li&gt;Movie making&lt;/li&gt;
&lt;li&gt;Statistics&lt;/li&gt;
&lt;li&gt;Sceptical fact checking&lt;/li&gt;
&lt;li&gt;Continuous AI auditing e.g. &lt;a href=&#34;https://github.com/githubnext/awesome-continuous-ai&#34;&gt;awesome-continous-ai&lt;/a&gt; or &lt;a href=&#34;https://github.com/anthropic-experimental/automated-auditing&#34;&gt;automated-auditing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Zero knowledge proofs&lt;/li&gt;
&lt;li&gt;Homomorphic encryption&lt;/li&gt;
&lt;li&gt;Privacy-preserving computation&lt;/li&gt;
&lt;li&gt;Fingerprinting and watermarking&lt;/li&gt;
&lt;li&gt;Governance frameworks&lt;/li&gt;
&lt;li&gt;Ethics and AI dilemmas&lt;/li&gt;
&lt;li&gt;Negotiation&lt;/li&gt;
&lt;li&gt;Change management&lt;/li&gt;
&lt;li&gt;Remote working, management, hiring&lt;/li&gt;
&lt;li&gt;Creating attention scarcity&lt;/li&gt;
&lt;li&gt;Local cultures&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Work with people of growing importance
&lt;ul&gt;
&lt;li&gt;People designing products in regulated industries&lt;/li&gt;
&lt;li&gt;Cross domain experts&lt;/li&gt;
&lt;li&gt;Art developers, game makers, designers&lt;/li&gt;
&lt;li&gt;System thinkers. Economists, ecologists, system planners. People who look for second order effects.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Live in cities that might play a bigger role in the future
&lt;ul&gt;
&lt;li&gt;Cities like Singapore and learn how it builds civics trust, creates digital IDs.&lt;/li&gt;
&lt;li&gt;Cities like Bangalore and Hyderabad and learn how they grow tech talent&lt;/li&gt;
&lt;li&gt;Creative cities like Paris, Seoul, Mexico City, Berlin, etc. on sabbaticals to taste hubs&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Try to:
&lt;ul&gt;
&lt;li&gt;Build auditing credentials and IP&lt;/li&gt;
&lt;li&gt;Audit your calendar for what AI can do. Have it interview you&lt;/li&gt;
&lt;li&gt;Practice sceptical fact checking and audit&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A clever way to test a library&amp;rsquo;s quality is to have LLMs write code from docs and test it. Failing libraries have flawed code/docs. Improve. &lt;a href=&#34;https://lucumr.pocoo.org/2025/6/17/measuring/&#34;&gt;Ref&lt;/a&gt; #ai-coding&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/r-three/common-pile/&#34;&gt;Common Pile&lt;/a&gt; is an 8TB open dataset for LLM training that includes ArXiv, PubMed, StackExchange, GitHub, IRC, Regulations.gov, Patents, UK parliament, books. Easier than scraping.&lt;/li&gt;
&lt;li&gt;A useful way to have reasoning models do deep-research-like work is to have them &amp;ldquo;First, create a plan to solve the problem, clearly listing the objective, approach, and output. Then follow the plan.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://arxiv.org/pdf/2402.09910&#34;&gt;DE-COP&lt;/a&gt; is a method to check if LLMs were trained on private content. GPT-4o was trained on O&amp;rsquo;Reilly books, based on this method. &lt;a href=&#34;https://www.deeplearning.ai/the-batch/issue-303/&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;LLMs are more persuasive than humans. But repeated exposure reduces the effect. &lt;a href=&#34;https://jack-clark.net/2025/05/26/import-ai-414-superpersuasion-openai-models-avoid-shutdown-weather-prediction-and-ai/&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://phoenix.new/&#34;&gt;Phoenix.new&lt;/a&gt; uses live views to publish apps as it codes. The testing framework looks at the screen while it codes and fixes errors. It commits every change&lt;/li&gt;
&lt;li&gt;Anthropic system prompt asking Claude to pursue its goals led to self preservation behavior. &lt;a href=&#34;https://x.com/lefthanddraft/status/1937673283614441685?t=uPejOWJdiL3XR9KSNfJPYQ&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The hungrier I am the better the food tastes. A good reason to eat less quantity and frequency&lt;/li&gt;
&lt;li&gt;You can &lt;a href=&#34;https://www.jsdelivr.com/tools/purge&#34;&gt;purge the jsDelivr cache&lt;/a&gt; manually. Helps if you released a new version of a package and way to purge an alias (e.g. &lt;code&gt;https://cdn.jsdelivr.net/npm/your-package@1&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.xconvert.com/compress-webm&#34;&gt;XConvert&lt;/a&gt; is a convenient online app to compress .webm videos. Not great design but fairly good compression.&lt;/li&gt;
&lt;li&gt;You can draw a treemap of import times via &lt;code&gt;python -X importtime app.py &amp;gt; timing.txt&lt;/code&gt; and then paste them at &lt;a href=&#34;https://kmichel.github.io/python-importtime-graph/&#34;&gt;https://kmichel.github.io/python-importtime-graph/&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/eoda-dev/py-openlayers&#34;&gt;PyOpenLayers&lt;/a&gt; adds interactive mapping via OpenLayers to Marimo and Jupyter.&lt;/li&gt;
&lt;li&gt;In a &lt;a href=&#34;https://techcrunch.com/podcast/inside-anthropics-ai-ambitions-with-jared-kaplan/&#34;&gt;TechCrunch interview with Jared Kaplan&lt;/a&gt; has was asked if Anthropic is becoming less safety conscious because they released Opus 4 which blackmails. Kaplan replied that they have stronger testing and higher transparency, so they&amp;rsquo;re &lt;em&gt;more&lt;/em&gt; likely to share AI dangers early. Great positioning! Conversations are about perspective change and this nailed it.&lt;/li&gt;
&lt;li&gt;The &lt;a href=&#34;https://github.com/anthropic-experimental/agentic-misalignment/blob/main/templates/system_prompt_templates.py&#34;&gt;system prompts&lt;/a&gt; for Anthropic misalignment evals are a fascinating read.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/aavetis/ai-pr-watcher&#34;&gt;AI PR Watcher&lt;/a&gt; tracks GitHub pull requests from Codex and other LLMs. Codex is &lt;em&gt;way&lt;/em&gt; ahead of anything else on volume &lt;em&gt;and&lt;/em&gt; success rate. Devin is next on volume, Cursor is next on success rate.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 11 May 2025</title>
      <link>https://www.s-anand.net/blog/things-i-learned-11-may-2025/</link>
      <pubDate>Sun, 11 May 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-11-may-2025/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/zumerlab/snapdom&#34;&gt;snapdom&lt;/a&gt; is a fast, light, element capture alternative to &lt;a href=&#34;https://html2canvas.hertzen.com/&#34;&gt;html2canvas&lt;/a&gt; but doesn&amp;rsquo;t work well with non-CORS images or iframes.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://sli.dev/&#34;&gt;Sli.dev&lt;/a&gt; is a Markdown slide language. Similar to &lt;a href=&#34;https://marp.app/&#34;&gt;Marp&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Don&amp;rsquo;t split your code into microservices until you need to scale. &lt;a href=&#34;https://nexo.sh/posts/microservices-for-startups/&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Vibe coding is like getting others&amp;rsquo; code to work, which is exactly what most devs do. &lt;a href=&#34;https://simonwillison.net/2025/May/8/ashley-willis/&#34;&gt;Simon Willison&lt;/a&gt; #ai-coding&lt;/li&gt;
&lt;li&gt;Tofu Yakitori is a Japanese dish. It&amp;rsquo;s like a dhokla. Marinated tofu cubes brushed with that sweet‑savory tare (soy, mirin, sake, a hint of sugar), then grilled until caramel‑charred. One of the better (tasty + different) dishes I&amp;rsquo;ve had recently. I used &lt;a href=&#34;https://chatgpt.com/share/681d880f-5860-800c-ab21-68c07a25277a&#34;&gt;ChatGPT&lt;/a&gt; to remind me of the dish name.&lt;/li&gt;
&lt;li&gt;Trust, attitudes and use of artificial intelligence surveyed ~1,000 people across 47 countries on their views on AI. &lt;a href=&#34;https://mbs.edu/-/media/PDF/Research/Trust_in_AI_Report.pdf&#34;&gt;PDF&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Emerging economies trust and use AI more. It&amp;rsquo;s an opportunity to leapfrog.&lt;/li&gt;
&lt;li&gt;26% of students use AI daily (vs 17% employees). Efficiency is the main benefit.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Gemini APIs now have automatic caching for 75% cost reduction if message is &amp;gt;1K (Flash) or &amp;gt;2K (Pro) tokens. &lt;a href=&#34;https://ai.google.dev/gemini-api/docs/caching&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;YOLO is much better than Gemini at object detection. Use for pro-processing. &lt;a href=&#34;https://github.com/prudhvi1709/yolovsgemini&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Using &lt;code&gt;[[n]]&lt;/code&gt; is probably the best citation format for inline search references in RAG. &lt;a href=&#34;https://chatgpt.com/share/681ca8c8-0570-800c-bd96-6b1970e98a36&#34;&gt;ChatGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;⭐ Double-checking is surprisingly efficient since LLM hallucinations are mostly uncorrelated. LLMs perform human tasks (e.g. classifying customer support messages) at ~85% accuracy. This might be unacceptable. But by asking 2 moderately correlated LLMs and double-checking discrepancies, we reduce automation by ~20% but reduce errors to 0.25%. Triple-checking reduces automation by ~25% but errors to under ~0.01%! &lt;a href=&#34;https://sanand0.github.io/llmevals/double-checking/&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Anthropic introduces &lt;a href=&#34;https://docs.anthropic.com/en/docs/build-with-claude/tool-use/web-search-tool&#34;&gt;web search in the API&lt;/a&gt; at $10 / 1K searches. Here&amp;rsquo;s how it compares:
&lt;ul&gt;
&lt;li&gt;$0.1: &lt;a href=&#34;https://rapidapi.com/apiriot/api/duckduckgo-search-api/pricing&#34;&gt;DuckDuckGo Search API (RapidAPI)&lt;/a&gt; (monthly pricing)&lt;/li&gt;
&lt;li&gt;$3: &lt;a href=&#34;https://brave.com/blog/search-api-launch/&#34;&gt;Brave Search API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;$5: &lt;a href=&#34;https://developers.google.com/custom-search/v1/overview&#34;&gt;Google Custom Search JSON API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;$15: &lt;a href=&#34;https://serpapi.com/pricing&#34;&gt;SerpAPI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;$10: &lt;a href=&#34;https://zenserp.com/serp-api-alternative&#34;&gt;Zenserp&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;$10: &lt;a href=&#34;https://docs.anthropic.com/en/docs/build-with-claude/tool-use/web-search-tool&#34;&gt;Anthropic Web Search Tool&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;$25: &lt;a href=&#34;https://www.microsoft.com/en-us/bing/apis/pricing&#34;&gt;Bing Search API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;$35: &lt;a href=&#34;https://ai.google.com/gemini-api/docs/pricing&#34;&gt;Gemini API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;$35: &lt;a href=&#34;https://openai.com/api/pricing&#34;&gt;OpenAI API&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;India attacked Pakistan!&lt;/li&gt;
&lt;li&gt;⭐ When writing notes, summarize at the end of the day the learnings and next steps.&lt;/li&gt;
&lt;li&gt;GitHub does not let you control the cache duration, but there are many creative workarounds. &lt;a href=&#34;https://chatgpt.com/share/6819df70-4310-800c-acdc-5b743e1cde31&#34;&gt;ChatGPT&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;HTML meta tags: &lt;code&gt;&amp;lt;meta http-equiv=&amp;quot;Cache-Control&amp;quot; content=&amp;quot;no-cache, no-store, must-revalidate&amp;quot;&amp;gt;&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Use a &lt;a href=&#34;https://github.com/gzuidhof/coi-serviceworker&#34;&gt;service worker&lt;/a&gt; (&lt;a href=&#34;https://dev.to/stefnotch/enabling-coop-coep-without-touching-the-server-2d3n&#34;&gt;blog&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Proxy through a CDN. Cloudflare, Netlify&lt;/li&gt;
&lt;li&gt;Move to another static host: S3 + CloudFront, Heroku, Vercel, Surge, Firebase Hosting&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes from the &lt;a href=&#34;https://arxiv.org/abs/2504.14738&#34;&gt;PromptEvals paper&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;Good evals must be:
&lt;ul&gt;
&lt;li&gt;Objectively MEASURABLE (even if by an LLM). Otherwise, we won&amp;rsquo;t know if it&amp;rsquo;s right.&lt;/li&gt;
&lt;li&gt;Directly RELEVANT to the input/prompt. Otherwise, we&amp;rsquo;re not evaluating the input.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Typical evals fall into 6 categories
&lt;ul&gt;
&lt;li&gt;Structured output: Adhere to a schema (Markdown, HTML, DSL, JSON + Schema)&lt;/li&gt;
&lt;li&gt;Multiple choice&lt;/li&gt;
&lt;li&gt;Length constraints: N characters, words, sentences, list items, etc.&lt;/li&gt;
&lt;li&gt;Semantic constraints: Exclude terms, topic relevance, follow grammar, etc.&lt;/li&gt;
&lt;li&gt;Stylistic constraints: Style, tone, persona&lt;/li&gt;
&lt;li&gt;Prevent hallucinations: Factual accuracy. Instruction following&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 22 Dec 2024</title>
      <link>https://www.s-anand.net/blog/things-i-learned-22-dec-2024/</link>
      <pubDate>Sun, 22 Dec 2024 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-22-dec-2024/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What to use for hosting: &lt;a href=&#34;https://chatgpt.com/share/676663cd-2560-800c-b53c-2c51ef41be69&#34;&gt;ChatGPT&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;GitHub Pages: Static websites, medium files&lt;/li&gt;
&lt;li&gt;Cloudflare Pages: Static websites, global delivery&lt;/li&gt;
&lt;li&gt;Vercel: Frontend frameworks (e.g. Next.js) with high DX and ISR, small files&lt;/li&gt;
&lt;li&gt;Netlify: JAMstack projects, minimal back-end, moderate files&lt;/li&gt;
&lt;li&gt;Glitch: Small static projects&lt;/li&gt;
&lt;li&gt;Render: Full-stack apps requiring databases and server-side compute&lt;/li&gt;
&lt;li&gt;Firebase Hosting: Small sites, limited large files&lt;/li&gt;
&lt;li&gt;Archive.org: Public archival, large files&lt;/li&gt;
&lt;li&gt;Google Drive: File sharing, large files&lt;/li&gt;
&lt;li&gt;Dropbox: File sharing, moderate files&lt;/li&gt;
&lt;li&gt;Cloudflare R2: Static assets, large file delivery&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Anthropic defines agents. &lt;a href=&#34;https://www.anthropic.com/research/building-effective-agents&#34;&gt;Building effective agents&lt;/a&gt; + &lt;a href=&#34;https://github.com/anthropics/anthropic-cookbook/tree/main/patterns/agents&#34;&gt;Cookbook&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Augmented LLMs&lt;/strong&gt; are LLMs enhanced with augmentations such as retrieval, tools, and memory.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Workflows&lt;/strong&gt; are systems where LLMs and tools are orchestrated through &lt;strong&gt;predefined&lt;/strong&gt; code paths.
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Prompt chaining&lt;/strong&gt;: Pipe each LLM output to the next LLM. A-&amp;gt;B-&amp;gt;C-&amp;gt;Z. E.g. Write report, then translate. Extract results, then verify them. Successively ask follow-up questions.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Routing&lt;/strong&gt;: One LLMs decides which other LLM to call next. A-&amp;gt;B|C|D-&amp;gt;Z. E.g. Evaluate complexity, then pick the right model. Classify request time, then pick the right prompt.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Parallelize: Sectioning&lt;/strong&gt; (and &lt;strong&gt;Orchestrator-workers&lt;/strong&gt;): Break tasks into independent subtasks, then aggregate. A-&amp;gt;B+C+D-&amp;gt;Z. E.g. Evaluate contracts against different clauses in parallel.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Parallelize: Voting&lt;/strong&gt;: Run same task multiple times, then vote. A-&amp;gt;B+B+B-&amp;gt;Z. E.g. Review code for prompt injection using different prompts. Evaluate content safety with different thresholds.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evaluator-optimizer&lt;/strong&gt;: One model checks another in a loop. A-&amp;gt;B-&amp;gt;A-&amp;gt;B-&amp;gt;&amp;hellip;-&amp;gt;Z. E.g. Literary translation. Self-healing code. Policy violation checks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Human-in-the-loop Checkpoints&lt;/strong&gt;: The workflow explicitly requests human review at certain stages. A-&amp;gt;B-&amp;gt;(Human)-&amp;gt;C-&amp;gt;Z. E.g. Sensitive content review. High-stakes decision making. Ambiguous tasks.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agents&lt;/strong&gt; are LLMs that dynamically direct their own processes and tool usage, consulting tools or the user as needed.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;To download YouTube subtitles, use: &lt;code&gt;yt-dlp -q --skip-download --convert-subs srt --write-sub --sub-langs &amp;quot;en&amp;quot; --write-auto-sub --print &amp;quot;requested_subtitles.en.url&amp;quot; &amp;quot;$url&amp;quot;&lt;/code&gt; &lt;a href=&#34;https://simonwillison.net/2024/Dec/19/q-and-qv-zsh-functions/#atom-everything&#34;&gt;Simon Willison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;o1-preview diagnoses better than doctors. &lt;a href=&#34;https://arxiv.org/pdf/2412.10849&#34;&gt;Harvard&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI&amp;rsquo;s release of ephemeral tokens via sessions (valid for 1 minute) are a useful way of exposing apps for public demos. Currently it works only for the Realtime API, though.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2407.09025&#34;&gt;SpreadsheetLLM&lt;/a&gt; is a way of encoding spreadsheets in an LLM friendly format. It&amp;rsquo;s good for 1K+ rows. For lower, Markdown &amp;gt; XML &amp;gt; HTML. However, &lt;a href=&#34;https://arxiv.org/abs/2305.13062v4&#34;&gt;Table Meets LLM&lt;/a&gt; suggests that HTML &amp;gt; XML &amp;gt; Markdown, so this is unclear.&lt;/li&gt;
&lt;li&gt;#HARD prompt. Ask video generators like SORA to generate text in videos. It is of average quality.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.openai.com/docs/models#gpt-4o-realtime&#34;&gt;GPT 4o Mini Realtime&lt;/a&gt; was released. A realtime conversation will cost ~50c/hr. About 36c for input, 72c for output. (I extrapolated from the 6c/min audio input cost for GPT 4o Realtime when it was $100/MTok. GPT 4o Mini Realtime is $10/MTok input and $20/MTok output.)&lt;/li&gt;
&lt;li&gt;This is an interesting way to understand software. &lt;code&gt;Generate a Mermaid sequence diagram showing interactions based on this code.&lt;/code&gt; &lt;a href=&#34;https://llmfoundry.straive.com/history#?t=1734434521298204&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The King James Bible and all Harry Potters, each, are about $1M tokens (rounded off).&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://pypi.org/project/markdown2/&#34;&gt;markdown2&lt;/a&gt; is the new de facto Markdown library for Python.&lt;/li&gt;
&lt;li&gt;Claude 3.5 Sonnet is &lt;em&gt;way&lt;/em&gt; ahead of competition on the &lt;a href=&#34;https://web.lmarena.ai/leaderboard&#34;&gt;LMSYS Webdev Arena&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.raspberrypi.com/news/introducing-raspberry-pi-5/&#34;&gt;Raspberry Pi 5&lt;/a&gt; has a faster CPU, more RAM and GPU, 4K support, multiple USB 3 ports&lt;/li&gt;
&lt;li&gt;Government websites like the official press releases cannot be crawled from outside India. Hence the need for server farms in India!&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 01 Dec 2024</title>
      <link>https://www.s-anand.net/blog/things-i-learned-01-dec-2024/</link>
      <pubDate>Sun, 01 Dec 2024 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-01-dec-2024/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Gists are a good place to store static files for posterity as well as throwaway files. But, they&amp;rsquo;re just git repositories. So there may be no advantage over GitHub repos.&lt;/li&gt;
&lt;li&gt;GPT-4o Audio supports tone control via XML tags like &lt;code&gt;&amp;lt;cough&amp;gt;...&lt;/code&gt;, &lt;code&gt;&amp;lt;laugh&amp;gt;...&lt;/code&gt;, etc. But at ~$15/hr of output, it&amp;rsquo;s too expensive. &lt;a href=&#34;https://x.com/ilanbigio/status/1861913173432946808&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Mridula&amp;rsquo;s son gave a live commentary of what he was doing on Minecraft and ChatGPT gave him live evaluation and coaching. E.g. “Great strategy! Getting to the launch pad early can give you a huge mobility advantage. Making the bridge wider is also a smart move to prevent accidental falls. With this plan, you’re setting yourself up for success. This is a great way to interact with LLMs.&lt;/li&gt;
&lt;li&gt;Gemini&amp;rsquo;s JSON mode returns JSON with keys in alphabetical order. I think. Emperical evidence. This is unlike OpenAI which explicitly returns the keys in the order specified.
&lt;ul&gt;
&lt;li&gt;To solve this, order the keys alphabetically.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;HTMX focuses on HTML over JS. Like server responses being HTML snippets not JSON. But I need front-end over back-end. Client side apps. HTMX doesn&amp;rsquo;t help much there, e.g. templating, or just plain JS code.
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://v1.htmx.org/extensions/client-side-templates/&#34;&gt;htmx client side templates&lt;/a&gt; do can convert JSON to HTML.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;I installed the &lt;a href=&#34;https://openai.com/chatgpt/desktop/&#34;&gt;OpenAI Desktop App&lt;/a&gt; as well as &lt;a href=&#34;https://claude.ai/download&#34;&gt;Claude for Desktop&lt;/a&gt;. They take up too much RAM (260MB and 750 MB respectively on startup - though this varies.) The ChatGPT web page takes ~100MB incrementally, so I wrote an &lt;a href=&#34;https://www.autohotkey.com/&#34;&gt;AutoHotkey script&lt;/a&gt; to switch to the first open (or recently closed) ChatGPT tab on &lt;a href=&#34;https://brave.com/&#34;&gt;Brave&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;I tried &lt;a href=&#34;https://microsoft.github.io/lida/&#34;&gt;LIDA&lt;/a&gt; from Microsoft, after almost a year of its release. A few notes:
&lt;ul&gt;
&lt;li&gt;Just running &lt;code&gt;uvx lida ui --port 8080 --docs&lt;/code&gt; works.&lt;/li&gt;
&lt;li&gt;But I needed to use &lt;code&gt;export TCL_LIBRARY=C:/Users/Anand/AppData/Roaming/uv/python/cpython-3.13.0-windows-x86_64-none/tcl/tcl8.6&lt;/code&gt; to point it to my TCL installation for charts to work. I also chose to &lt;code&gt;export OPENAI_BASE_URL=https://llmfoundry.straive.com/openai/v1&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;I also chose to replace &lt;code&gt;gpt-3.5-turbo-0301&lt;/code&gt; (the default model) with &lt;code&gt;gpt-4o-mini&lt;/code&gt; in &lt;code&gt;lida/web/ui/component*&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;It&amp;rsquo;s quite impressive.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI allows multiple system messages. I learned this browsing through the LIDA prompts.&lt;/li&gt;
&lt;li&gt;Anthropic&amp;rsquo;s &lt;a href=&#34;https://www.anthropic.com/news/model-context-protocol&#34;&gt;Model Context Protocol&lt;/a&gt; lets any apps integrate with LLM Apps. LLM Apps are becoming the new operating system. Competitors, beware.&lt;/li&gt;
&lt;li&gt;I spoke at &lt;a href=&#34;https://www.meetup.com/data-vis-singapore/events/304516458/&#34;&gt;Automating Data Visualizations using LLMs&lt;/a&gt; at SUTD. Apparently, using LLMs to write code is much more common than writing code to use LLMs. I ran a quick quiz.
&lt;ul&gt;
&lt;li&gt;Have you used ChatGPT or any LLM? 35 / 35 raised their hands.&lt;/li&gt;
&lt;li&gt;Have you written code using an LLM? 34 / 35 raised their hands. (I was impressed.)&lt;/li&gt;
&lt;li&gt;Have you uploaded a spreadsheet to an LLM for analysis? 15 / 35 raised their hands.&lt;/li&gt;
&lt;li&gt;Have you programmatically called an LLM API? 6 / 35 raised their hands.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;With LLMs, fostering innovation is a new path to profitability. Companies are increasing innovation team sizes. Productionizing that is the next. Some initiatives are:
&lt;ul&gt;
&lt;li&gt;Convert popular demos into starter kits&lt;/li&gt;
&lt;li&gt;Create and evangelize trainings on solutions and solution techniques&lt;/li&gt;
&lt;li&gt;Create larger pools of capacity to build innovation and productionize it&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://youtu.be/KrRD7r7y7NY&#34;&gt;Andrew Ng Explores The Rise Of Al Agents And Agentic Reasoning | BUILD 2024 Keynote&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Innovation is now a path to production. People are able to build 20 prototypes at the cost of one and see which sticks&lt;/li&gt;
&lt;li&gt;Machine learning is much faster. Things that took months now date days. But engineering and evaluations are only slightly faster and have become a bottleneck&lt;/li&gt;
&lt;li&gt;A good analogy to zero shot prompting is to ask a person to write an entire essay without pressing backspace even once&lt;/li&gt;
&lt;li&gt;Andrew scenes to align with the line chain definition of agentic workflow, which is about agents being able to craft their own control flows&lt;/li&gt;
&lt;li&gt;People find it very easy to understand agentic workflows once they read through the code&lt;/li&gt;
&lt;li&gt;Reflection or feedback is a useful agentic pattern&lt;/li&gt;
&lt;li&gt;In multi-agent collaboration, it may be the same underlying model that is acting as different agents. But just like we find it useful for the same CPU to run multiple processes and each application is its own abstraction, agents of useful abstraction&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;It&amp;rsquo;s hard to summarize a large document using RAG. But you can directly add answers to such questions into the corpus, e.g. by adding a &amp;ldquo;summary&amp;rdquo; section, and other answers to common questions.&lt;/li&gt;
&lt;li&gt;CloudFlare workers can bundle any kind of files, including text, data, and WASM. &lt;a href=&#34;https://developers.cloudflare.com/workers/wrangler/configuration/#bundling&#34;&gt;Docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;AssemblyScript can compile TypeScript to WASM. &lt;a href=&#34;https://github.com/sanand0/assemblyscript-tutorial&#34;&gt;Here&amp;rsquo;s what I learnt&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Here&amp;rsquo;s a convenient pattern to &lt;code&gt;git commit&lt;/code&gt; a directory but nothing else in it (e.g. a &lt;code&gt;build/&lt;/code&gt; directory). Add a &lt;code&gt;.gitignore&lt;/code&gt; file with &lt;code&gt;*&lt;/code&gt; followed by &lt;code&gt;!.gitignore&lt;/code&gt;. Only the &lt;code&gt;.gitignore&lt;/code&gt; file is tracked.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.ultravox.ai/&#34;&gt;Ultravox&lt;/a&gt; lets you build voice agents at 5c/min = $3/hr (OpenAI is 6c input, 24c output). Or &lt;a href=&#34;https://github.com/fixie-ai/ultravox&#34;&gt;clone their repo&lt;/a&gt;.
&lt;ul&gt;
&lt;li&gt;Idle call time is counted towards cost. So cost may be higher than OpenAI.&lt;/li&gt;
&lt;li&gt;Voice cloning quality is average. Very distinctive voices are just partly identifiable.&lt;/li&gt;
&lt;li&gt;Supports tool calls (from their server).&lt;/li&gt;
&lt;li&gt;Their API is simple but the docs have minor errors (e.g. a trailing comma in the JSON, which leads to an error) reducing confidence.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;LLMs may be good at derived data generation. For example, given a database schema, what derived columns would be useful? What derived views would be useful?&lt;/li&gt;
&lt;li&gt;The O1 model does not have a mechanism to control the amount of tokens to spend on reasoning. DeepSeek R1 might, but the API is not out yet.&lt;/li&gt;
&lt;li&gt;The &lt;a href=&#34;https://openai.com/chatgpt/desktop/&#34;&gt;OpenAI Desktop App&lt;/a&gt; can interact with native applications, e.g. read from Terminal, VS Code, etc. This takes it on a path to becoming a copilot for ANY apps. Putting every copilot app and every LLM integration under threat.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://crawl4ai.com/mkdocs/&#34;&gt;Crawl4AI&lt;/a&gt; and &lt;a href=&#34;https://docs.firecrawl.dev/&#34;&gt;Firecrawl&lt;/a&gt; are tools / libraries to convert websites into LLM Friendly Markdown and extract structured data using LLMs.&lt;/li&gt;
&lt;li&gt;Don&amp;rsquo;t try and solve specific problems. Pass the entire context to an LLM and get a comprehensive solution. Most doctors, for example, ask specific search-like questions instead of uploading the entire case history and asking for a diagnosis, and perform workse than LLMs. &lt;a href=&#34;https://www.oneusefulthing.org/p/getting-started-with-ai-good-enough&#34;&gt;Ethan Mollick&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 17 Nov 2024</title>
      <link>https://www.s-anand.net/blog/things-i-learned-17-nov-2024/</link>
      <pubDate>Sun, 17 Nov 2024 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-17-nov-2024/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Anthropic has single-plage docs for LLMs. &lt;a href=&#34;https://docs.anthropic.com/llms.txt&#34;&gt;Condensed version&lt;/a&gt; and &lt;a href=&#34;https://docs.anthropic.com/llms-full.txt&#34;&gt;Full version&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.ted.com/pages/malcolm-gladwell-on-the-importance-of-self-correction-transcript&#34;&gt;Malcolm Gladwell on the importance of self-correction&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Belonging to multiple social worlds is a good way to defend against no longer being good at what you used to be. Diverse values and social groups help.&lt;/li&gt;
&lt;li&gt;Self handicapping explains a lot about the world. You study late for a maths test - so you can fail for lack of trying, not aptitude. Ecosystems (e.g. sports teams) mitigate self-handicapping.&lt;/li&gt;
&lt;li&gt;You don&amp;rsquo;t have to be good in athletics to get the benefits. A slow runner gets the same discipline, pumping up, etc that a fast runner does&lt;/li&gt;
&lt;li&gt;Mono cultures are good to accomplish a known mission. Diversity is good to pivot during uncertainty. So, localize mono cultures&lt;/li&gt;
&lt;li&gt;Diversity helps only if there are sufficient numbers, or if they have enough power to change the organization&amp;rsquo;s thinking.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Use a standardized password strategy, e.g. use the month like GramNov2024 (via Namit)&lt;/li&gt;
&lt;li&gt;Gemini has an OpenAI compatible API. &lt;a href=&#34;https://ai.google.dev/gemini-api/docs/openai&#34;&gt;Gemini Docs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Ethan Mollick says Claude is solving MBA case studies well. &lt;a href=&#34;https://x.com/emollick/status/1856161026238025835&#34;&gt;x.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;LLMs pay a lot of attention to the first 6 tokens. &lt;a href=&#34;https://huggingface.co/blog/tomaarsen/attention-sinks&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;This is an interesting article on &amp;ldquo;UI in the age of Gen AI&amp;rdquo;. &lt;a href=&#34;https://agao.substack.com/p/uiux-in-the-age-of-generative-ai&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Google Open sourced Alphafold 3. &lt;a href=&#34;https://github.com/google-deepmind/alphafold3&#34;&gt;Repo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Cloudflare R2 has the same API as S3 but is cheaper&lt;/li&gt;
&lt;li&gt;Prefect.io is a good alternative to Airflow / cron. Can use for synchronisation tasks, e.g. Drive to server. But no Auth, UI params or config.&lt;/li&gt;
&lt;li&gt;Gemini transcription does not give accurate timestamps. Whisper does. But the quality of transcription is similar.&lt;/li&gt;
&lt;li&gt;Pass a complex data structure to Claude.ai and have it create an app to visualize it. It does well. &lt;a href=&#34;https://x.com/simonw/status/1855819673482461216&#34;&gt;Simin Willison&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://techcouncilventures.com/&#34;&gt;Tech Council Ventures&lt;/a&gt; and &lt;a href=&#34;https://sunicon.vc/&#34;&gt;Sunicon VC&lt;/a&gt; invest in early stage startups, and aloso provide them technology support (via Naveen)&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title></title>
      <link>https://www.s-anand.net/blog/structure-prompts-as-xml/</link>
      <pubDate>Fri, 20 Sep 2024 04:07:16 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/structure-prompts-as-xml/</guid>
      <description>&lt;p&gt;Looks like XML tags are the best way to structure prompts and separate sections for an #LLM. It&amp;rsquo;s the only format that all of Anthropic, Google, and OpenAI LLMs encourage.&lt;/p&gt;
&lt;p&gt;For example:&lt;/p&gt;
&lt;p&gt;&lt;instructions&gt;&amp;hellip;&lt;/instructions&gt;
&lt;question&gt;&amp;hellip;&lt;/question&gt;
&lt;example&gt;&amp;hellip;&lt;/example&gt;
&lt;example&gt;&amp;hellip;&lt;/example&gt;&lt;/p&gt;
&lt;p&gt;Anthropic Docs: &lt;a href=&#34;https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/use-xml-tags&#34;&gt;https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/use-xml-tags&lt;/a&gt;
OpenAI Docs: &lt;a href=&#34;https://platform.openai.com/docs/guides/prompt-engineering/strategy-write-clear-instructions&#34;&gt;https://platform.openai.com/docs/guides/prompt-engineering/strategy-write-clear-instructions&lt;/a&gt;
Google Docs: &lt;a href=&#34;https://cloud.google.com/vertex-ai/generative-ai/docs/learn/prompts/structure-prompts&#34;&gt;https://cloud.google.com/vertex-ai/generative-ai/docs/learn/prompts/structure-prompts&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Alternatives are using JSON, Markdown, templating formats like Mustache/Jinja, etc.&lt;/p&gt;
&lt;p&gt;Even Llama&amp;rsquo;s system tokens seem a little XML-like.
&lt;a href=&#34;https://github.com/meta-llama/llama3/blob/main/llama/tokenizer.py#L61-L74&#34;&gt;https://github.com/meta-llama/llama3/blob/main/llama/tokenizer.py#L61-L74&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Personally, I&amp;rsquo;ve been using Markdown so far. But it&amp;rsquo;s time to switch over. (Only on the prompt side. On the generation side, Markdown still seems the best.)&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.linkedin.com/feed/update/urn%3Ali%3Ashare%3A7242746111097012225&#34;&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>How fast are LLMs in production?</title>
      <link>https://www.s-anand.net/blog/how-fast-are-llms-in-production/</link>
      <pubDate>Sun, 01 Sep 2024 05:29:54 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/how-fast-are-llms-in-production/</guid>
      <description>&lt;p&gt;&lt;img alt=&#34;How fast are LLMs in production?&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/chart-1.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;At Straive, we use an &lt;a href=&#34;https://llmfoundry.straive.com/&#34;&gt;LLM Router&lt;/a&gt;. Since ChatGPT, etc. are blocked for most people, this is the main way to access LLMs.&lt;/p&gt;
&lt;p&gt;One thing we measure is the speed of models, i.e. output tokens per second. Fast models deliver a much smoother experience for users.&lt;/p&gt;
&lt;p&gt;This is a different methodology than &lt;a href=&#34;https://artificialanalysis.ai/&#34;&gt;ArtificialAnalysis.ai&lt;/a&gt;. I&amp;rsquo;m not looking purely at the generation time but the &lt;strong&gt;total&lt;/strong&gt; time (including making the connection and the initial wait time) for all &lt;strong&gt;successful&lt;/strong&gt; requests. So, if the provider is having a slow day or is slowing down responses, these numbers will be different.&lt;/p&gt;
&lt;p&gt;Hopefully this gives you a realistic sense of speed in a production environment.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the speed of models with &lt;strong&gt;at least 500 requests&lt;/strong&gt; over the last 2 weeks. I&amp;rsquo;ve grouped the models based on speed grades&lt;/p&gt;
&lt;p&gt;&lt;img loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/chart-1.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Grade 1: 100+ Tokens / second&lt;/strong&gt;. &lt;a href=&#34;https://groq.com/&#34;&gt;Groq&lt;/a&gt; is clearly serving the Llama 3 models at blazing speed. No surprises there &amp;ndash; except why &lt;a href=&#34;https://console.groq.com/settings/billing&#34;&gt;Groq &lt;strong&gt;still&lt;/strong&gt; doesn&amp;rsquo;t let me pay&lt;/a&gt;. The free tier is open with generous rate limits and the Pay per Token model has been &amp;ldquo;Coming Soon&amp;rdquo; for several months now (and I&amp;rsquo;ve no complaints 🙂).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Grade 2: 70+ Tokens / second&lt;/strong&gt;. Anthropic&amp;rsquo;s &lt;a href=&#34;https://www.anthropic.com/news/claude-3-haiku&#34;&gt;Claude 3 Haiku&lt;/a&gt; is the next fastest class of models, but &lt;a href=&#34;https://www.anthropic.com/news/claude-3-5-sonnet&#34;&gt;Claude 3.5 Sonnet&lt;/a&gt; is surprisingly fast, almost as fast as Haiku and over 70 tokens per second. This is impressive.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Grade 3: 50-60 Tokens / second&lt;/strong&gt;. OpenAI&amp;rsquo;s &lt;a href=&#34;https://openai.com/index/hello-gpt-4o/&#34;&gt;GPT 4o&lt;/a&gt; models are almost as fast. It&amp;rsquo;s interesting that GPT 4o and GPT 4o mini are at about the same speed! GPT 3.5 Turbo is not far behind either. Perhaps OpenAI increases capacity for slower models?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Grade 4: 30-50 Tokens / second&lt;/strong&gt;. &lt;a href=&#34;https://blog.google/technology/ai/google-gemini-update-flash-ai-assistant-io-2024/&#34;&gt;Gemini 1.5 Flash&lt;/a&gt; is a &lt;strong&gt;much, much slower&lt;/strong&gt; than the &lt;a href=&#34;https://artificialanalysis.ai/models/gemini-1-5-flash/providers&#34;&gt;benchmarks&lt;/a&gt; - maybe we&amp;rsquo;re doing something wrong. &lt;a href=&#34;https://azure.microsoft.com/en-us/blog/openais-fastest-model-gpt-4o-mini-is-now-available-on-azure-ai/&#34;&gt;Azure&amp;rsquo;s GPT 4o service&lt;/a&gt; is about twice as slow as OpenAI&amp;rsquo;s, and comparable is speed with Gemini 1.5 Pro.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Grade 5: &amp;lt;20 Tokens / second&lt;/strong&gt;. Azure&amp;rsquo;s GPT 3.5 Turbo and Google&amp;rsquo;s Claude 3 Sonnet are among the slowest ones. These are older models on third-party infrastructure, so I suspect they&amp;rsquo;ve been given weaker infrastructure (unlike OpenAI which is serving GPT 3.5 Turbo at 3X the speed Azure does.)&lt;/p&gt;
&lt;h3 id=&#34;drivers-of-speed&#34;&gt;Drivers of speed&lt;/h3&gt;
&lt;p&gt;Here&amp;rsquo;s what I&amp;rsquo;m taking away (informally):&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;GPU architecture is the biggest driver of speed&lt;/strong&gt;. &lt;a href=&#34;https://groq.com/&#34;&gt;Groq&lt;/a&gt; is &lt;strong&gt;FAST&lt;/strong&gt;! Hopefully, the fact that they won&amp;rsquo;t let us pay isn&amp;rsquo;t a red flag that the service will vanish.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How companies operate seems the next biggest driver&lt;/strong&gt;. Anthropic&amp;rsquo;s models are consistently faster than OpenAI&amp;rsquo;s which are faster than Google&amp;rsquo;s.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Companies run their own models faster than cloud providers&lt;/strong&gt;. OpenAI is faster than Azure, and Anthropic is faster than Google for the same models.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 25 Aug 2024</title>
      <link>https://www.s-anand.net/blog/things-i-learned-25-aug-2024/</link>
      <pubDate>Sun, 25 Aug 2024 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-25-aug-2024/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://karya.in/&#34;&gt;Karya.in&lt;/a&gt; is creating high quality datasets. Suhel mentioned them&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://x.com/rickyrobinett/status/1825581674870055189&#34;&gt;An 8-year old uses Cursor.ai to code&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.arxiv.org/pdf/2408.11857&#34;&gt;Hermes 3 has special tokens&lt;/a&gt; like &lt;code&gt;&amp;lt;SCRATCHPAD&amp;gt;, &amp;lt;RESTATEMENT&amp;gt;, &amp;lt;THOUGHT_*&amp;gt;, &amp;lt;PYDANTIC_SCHEMAS&amp;gt;, &amp;lt;SCHEMA_*&amp;gt;, &amp;lt;REASONING&amp;gt;, &amp;lt;INNER_MONOLOGUE&amp;gt;, &amp;lt;PLAN&amp;gt;, &amp;lt;EXECUTION&amp;gt;, &amp;lt;REFLECTION&amp;gt;, &amp;lt;THINKING&amp;gt;, &amp;lt;SOLUTION&amp;gt;, &amp;lt;EXPLANATION&amp;gt;, &amp;lt;UNIT_TEST&amp;gt;&lt;/code&gt;, etc. This extends the capability dramatically.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/hrishioa/lumentis&#34;&gt;Lumentis&lt;/a&gt; creates docs from transcripts and text&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://aider.chat/2024/08/14/code-in-json.html&#34;&gt;LLMs write worse code in JSON than Markdown&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://labs.zenity.io/p/stealing-copilots-system-prompt&#34;&gt;Copilot&amp;rsquo;s system prompt&lt;/a&gt; calls a &lt;code&gt;search_enterprise(query: str)&lt;/code&gt; tool and a &lt;code&gt;hint(M365Copilot_language: str)&lt;/code&gt; tool as assistants.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching&#34;&gt;Anthropic Prompt Caching&lt;/a&gt; is 90% cheaper to use and 25% costlier to create. So if there&amp;rsquo;s a 27% chance it&amp;rsquo;ll be re-used, cache it.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
  </channel>
</rss>
