<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>voice-cloning on S Anand</title>
    <link>https://www.s-anand.net/blog/tag/voice-cloning/</link>
    <description>Recent content in voice-cloning on S Anand</description>
    <generator>Hugo -- 0.164.0</generator>
    <language>en-us</language>
    <lastBuildDate>Sun, 29 Mar 2026 22:31:33 +0800</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/tag/voice-cloning/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>MGR via ElevenLabs</title>
      <link>https://www.s-anand.net/blog/mgr-via-elevenlabs/</link>
      <pubDate>Sun, 29 Mar 2026 22:31:33 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/mgr-via-elevenlabs/</guid>
      <description>&lt;p&gt;I was watching &lt;a href=&#34;https://en.wikipedia.org/wiki/Vaa_Vaathiyaar&#34;&gt;Vaa Vaathiyar&lt;/a&gt; which has a short clip of &lt;a href=&#34;https://en.wikipedia.org/wiki/M._G._Ramachandran&#34;&gt;MGR&lt;/a&gt; speaking. It&amp;rsquo;s either AI-generated or mimic-ed and it wasn&amp;rsquo;t bad.&lt;/p&gt;
&lt;p&gt;&lt;audio controls src=&#34;https://files.s-anand.net/images/2026-03-29-mgr-in-film.opus&#34;&gt;&lt;/audio&gt;&lt;/p&gt;
&lt;p&gt;I used &lt;code&gt;ffmpeg&lt;/code&gt; to record the audio from the film, transcribed it via &lt;a href=&#34;https://aistudio.google.com/prompts/new_chat?model=gemini-3.1-pro-preview&#34;&gt;Gemini 3 Pro on AI Studio&lt;/a&gt; with the prompt:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Transcribe this into Tamil&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&amp;hellip; which gave me:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;ராமு&amp;hellip;
என்ன செய்திருக்கிறாய் நீ&amp;hellip;
வாத்தியார் கேட்கிறேன் சொல்
நிமிர்ந்து பார்க்க கூட தைரியம் இல்லையா&amp;hellip;
ஓடாதே&amp;hellip; நில்&amp;hellip;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Translation:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Ramu&amp;hellip;
What have you done&amp;hellip;
Vaathiyar (MGR) is asking, tell me
Don&amp;rsquo;t you have the courage to stand up and look at me&amp;hellip;
Don&amp;rsquo;t run&amp;hellip; stop&amp;hellip;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(GitHub Copilot&amp;rsquo;s auto-complete translated the above for me as I typed - flawlessly. It&amp;rsquo;s getting better by the day!)&lt;/p&gt;
&lt;p&gt;Then, I used &lt;a href=&#34;https://github.com/yt-dlp/yt-dlp&#34;&gt;yt-dlp&lt;/a&gt; to download the audio from this &lt;a href=&#34;https://www.youtube.com/shorts/1jQqKds2z7g&#34;&gt;MGR Short Clip&lt;/a&gt;.&lt;/p&gt;
&lt;iframe width=&#34;560&#34; height=&#34;315&#34; src=&#34;https://www.youtube.com/embed/1jQqKds2z7g?si=8GUnT7IvcsQZpwDb&#34; title=&#34;YouTube video player&#34; frameborder=&#34;0&#34; allow=&#34;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&#34; referrerpolicy=&#34;strict-origin-when-cross-origin&#34; allowfullscreen&gt;&lt;/iframe&gt;
&lt;p&gt;Here&amp;rsquo;s the sample:&lt;/p&gt;
&lt;p&gt;&lt;audio controls src=&#34;https://files.s-anand.net/images/2026-03-29-mgr-sample.opus&#34;&gt;&lt;/audio&gt;&lt;/p&gt;
&lt;p&gt;I fed this into ElevenLabs&amp;rsquo; &lt;a href=&#34;https://elevenlabs.io/app/voice-library&#34;&gt;Instant Voice Clone&lt;/a&gt; that needs just 10 seconds of audio and created an &amp;ldquo;MGR&amp;rdquo; voice.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the same dialogue in the cloned voice:&lt;/p&gt;
&lt;p&gt;&lt;audio controls src=&#34;https://files.s-anand.net/images/2026-03-29-mgr-generated.opus&#34;&gt;&lt;/audio&gt;&lt;/p&gt;
&lt;p&gt;Personally, I think the ElevenLabs version is &lt;em&gt;slightly&lt;/em&gt; better. Of course, given the pace of AI improvement, this might just be the impact of a new model release.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;2 Apr 2026&lt;/strong&gt;: Here&amp;rsquo;s the non-cloned generation from &lt;a href=&#34;https://dashboard.sarvam.ai/text-to-speech&#34;&gt;Sarvam&amp;rsquo;s text to speech&lt;/a&gt; with Bulbul v3 standard quality. It feels pretty weak.&lt;/p&gt;
&lt;p&gt;&lt;audio controls src=&#34;https://files.s-anand.net/images/2026-03-29-mgr-sarvam.opus&#34;&gt;&lt;/audio&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://aistudio.google.com/u/2/generate-speech?model=gemini-2.5-pro-preview-tts&#34;&gt;Gemini 2.5 Pro Preview TTS&lt;/a&gt; gave me this, which feels &lt;em&gt;much&lt;/em&gt; better.&lt;/p&gt;
&lt;p&gt;&lt;audio controls src=&#34;https://files.s-anand.net/images/2026-03-29-mgr-gemini.opus&#34;&gt;&lt;/audio&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Voice coding is the new live coding</title>
      <link>https://www.s-anand.net/blog/voice-coding-is-the-new-live-coding/</link>
      <pubDate>Sun, 21 Sep 2025 11:18:17 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/voice-coding-is-the-new-live-coding/</guid>
      <description>&lt;p&gt;&lt;img alt=&#34;Voice coding is the new live coding&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/ChatGPT-Image-Sep-21-2025-04_46_27-PM.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;In Feb 2025 at PyConf Hyderabad, I tried a new slide format: &lt;a href=&#34;https://www.s-anand.net/blog/command-line-slideshows-in-bash/&#34;&gt;command-line slideshows in &lt;code&gt;bash&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve used this format in more talks since then:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/talks0/blob/main/2025-06-pycon-sg/llm-cli.md&#34;&gt;LLMs in the CLI&lt;/a&gt;, PyCon Singapore, Jun 2025&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/talks/blob/main/2025-07-24-pugs-agent-loop/README.md&#34;&gt;Agents in the CLI&lt;/a&gt;, Singapore Python User Group, Jul 2025&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sanand0/talks/blob/main/2025-09-13-duckdb-is-the-new-pandas/README.md&#34;&gt;DuckDB is the new Pandas&lt;/a&gt;, PyCon India, Sep 2025&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It&amp;rsquo;s my favorite format. I can demo code without breaking the presentation flow.&lt;br&gt;
It also draws interest. My setup was the &lt;a href=&#34;https://github.com/sanand0/talks/blob/main/2025-09-13-duckdb-is-the-new-pandas/README.md#qa&#34;&gt;top question in my PyCon talk&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;In Sep 2025, at PyCon India, I extended the setup for voice typing. &lt;a href=&#34;https://github.com/sanand0/scripts/blob/54560718bf2f4148d9005d74ab1543de52cff6d9/talkcode.sh&#34;&gt;&lt;code&gt;talkcode.sh&lt;/code&gt;&lt;/a&gt; is a Bash pipeline that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;uses &lt;a href=&#34;https://ffmpeg.org/&#34;&gt;&lt;code&gt;ffmpeg&lt;/code&gt;&lt;/a&gt; to record mic into as 16 kHz mono, voice-optimized &lt;code&gt;.opus&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;sends audio to Gemini via &lt;a href=&#34;https://llm.datasette.io/&#34;&gt;&lt;code&gt;llm&lt;/code&gt;&lt;/a&gt; for transcription&lt;/li&gt;
&lt;li&gt;uses &lt;a href=&#34;https://en.wikipedia.org/wiki/AWK&#34;&gt;&lt;code&gt;awk&lt;/code&gt;&lt;/a&gt; to extract the code fence&lt;/li&gt;
&lt;li&gt;uses &lt;a href=&#34;https://github.com/astrand/xclip&#34;&gt;&lt;code&gt;xclip&lt;/code&gt;&lt;/a&gt; to copy the code to the clipboard&lt;/li&gt;
&lt;li&gt;uses &lt;a href=&#34;https://github.com/jordansissel/xdotool&#34;&gt;&lt;code&gt;xdotool&lt;/code&gt;&lt;/a&gt; to paste it back to the original window&lt;/li&gt;
&lt;/ul&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-bash&#34; data-lang=&#34;bash&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Mic only, 16 kHz mono, voice filtering, fast Opus&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;ffmpeg -hide_banner -v error &lt;span class=&#34;se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  -f pulse -i default &lt;span class=&#34;se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  -ac &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt; -ar &lt;span class=&#34;m&#34;&gt;16000&lt;/span&gt; &lt;span class=&#34;se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  -af &lt;span class=&#34;s2&#34;&gt;&amp;#34;highpass=f=100,lowpass=f=6000&amp;#34;&lt;/span&gt; &lt;span class=&#34;se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  -c:a libopus -b:a 16k -vbr on &lt;span class=&#34;se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  -compression_level &lt;span class=&#34;m&#34;&gt;2&lt;/span&gt; -application voip -frame_duration &lt;span class=&#34;m&#34;&gt;60&lt;/span&gt; &lt;span class=&#34;se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  -y &lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt;&lt;span class=&#34;nv&#34;&gt;$AUDIO&lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Transcribe, extract the code fence and copy to clipboard&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;llm -m gemini-2.5-flash -a &lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt;&lt;span class=&#34;nv&#34;&gt;$AUDIO&lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt; -s &lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt;&lt;span class=&#34;nv&#34;&gt;$SYSTEM_TEXT&lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt; &lt;span class=&#34;se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;p&#34;&gt;|&lt;/span&gt; tee /dev/tty &lt;span class=&#34;se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;p&#34;&gt;|&lt;/span&gt; awk &lt;span class=&#34;s1&#34;&gt;&amp;#39;BEGIN{f=0} /```/{f=!f; next} f{buf=buf$0&amp;#34;\n&amp;#34;} END{print buf}&amp;#39;&lt;/span&gt; &lt;span class=&#34;se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;p&#34;&gt;|&lt;/span&gt; xclip -selection clipboard
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Bring last window to foreground and paste from clipboard&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;if&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;[&lt;/span&gt; -n &lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt;&lt;span class=&#34;si&#34;&gt;${&lt;/span&gt;&lt;span class=&#34;nv&#34;&gt;ACTIVE_WIN&lt;/span&gt;&lt;span class=&#34;k&#34;&gt;:-&lt;/span&gt;&lt;span class=&#34;si&#34;&gt;}&lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;]&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;;&lt;/span&gt; &lt;span class=&#34;k&#34;&gt;then&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  xdotool windowactivate --sync &lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt;&lt;span class=&#34;nv&#34;&gt;$ACTIVE_WIN&lt;/span&gt;&lt;span class=&#34;s2&#34;&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  sleep 0.08
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  xdotool key --clearmodifiers ctrl+shift+v
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;k&#34;&gt;fi&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;During the workshop, I said, &amp;ldquo;Which (judicial) bench has the longest pending cases,&amp;rdquo; and it generated a DuckDB query that, single-shot, ran correctly on the &lt;a href=&#34;https://github.com/vanga/indian-high-court-judgments&#34;&gt;Indian High Court judgements&lt;/a&gt; dataset.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;But LLMs are slow and break the flow. Here&amp;rsquo;s how I keep the room engaged:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Dictate, don&amp;rsquo;t type.&lt;/strong&gt; Speaking is faster &lt;strong&gt;and&lt;/strong&gt; more engaging. That&amp;rsquo;s why I built this workflow.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Avoid Alt-Tab.&lt;/strong&gt; Bring the LLM &lt;strong&gt;into&lt;/strong&gt; your app. Window-switching and copy-paste break focus for you &lt;strong&gt;and&lt;/strong&gt; the audience.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Always stream output.&lt;/strong&gt; Narrate as it loads. &lt;code&gt;| tee /dev/tty&lt;/code&gt; streams while piping.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Answer questions while you wait.&lt;/strong&gt; Keep a &lt;a href=&#34;https://www.slido.com/&#34;&gt;Slido&lt;/a&gt; Q&amp;amp;A open and address the top ones while waiting for the LLM.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;a href=&#34;https://www.linkedin.com/posts/sanand0_%F0%9D%97%A9%F0%9D%97%BC%F0%9D%97%B6%F0%9D%97%B0%F0%9D%97%B2-%F0%9D%97%B0%F0%9D%97%BC%F0%9D%97%B1%F0%9D%97%B6%F0%9D%97%BB%F0%9D%97%B4-%F0%9D%97%B6%F0%9D%98%80-%F0%9D%98%81%F0%9D%97%B5%F0%9D%97%B2-%F0%9D%97%BB%F0%9D%97%B2%F0%9D%98%84-activity-7376824026595278848-PxgV&#34;&gt;LinkedIn&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 03 Nov 2024</title>
      <link>https://www.s-anand.net/blog/things-i-learned-03-nov-2024/</link>
      <pubDate>Sun, 03 Nov 2024 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-03-nov-2024/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Indian companies with 30+ employees MUST have 2.5%-15% of their employees as apprentices. &lt;a href=&#34;https://chatgpt.com/share/67249206-3be0-800c-a2c3-d1dab166f180&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.textnow.com/&#34;&gt;Textnow&lt;/a&gt; and &lt;a href=&#34;https://textfree.us/&#34;&gt;TextFree&lt;/a&gt; provides a free phone number (like a virtual SIM). (But TextFree has more ads.) Keep using to avoid deactivation. No guarantee of retaining the number.
&lt;ul&gt;
&lt;li&gt;Some banks don&amp;rsquo;t accept TextNow for verification SMS. But voice call is OK.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.tello.com/&#34;&gt;Tello&lt;/a&gt;, &lt;a href=&#34;https://www.redpocket.com/&#34;&gt;Red pocket&lt;/a&gt; are cheap MVNOs with $5/month voice plans.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.metrobyt-mobile.com/&#34;&gt;Metro by T-Mobile&lt;/a&gt; and &lt;a href=&#34;cricketwireless.com&#34;&gt;Cricket&lt;/a&gt; are other MVNOs.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://mintmobile.com/&#34;&gt;MintMobile&lt;/a&gt; and &lt;a href=&#34;https://usmobile.com/&#34;&gt;US Mobile&lt;/a&gt; have $15/month and $8/month data plans.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The scientific discoveries that might have remained undiscovered for long if not for their discoverers &lt;a href=&#34;https://chatgpt.com/share/6722ec5e-56b8-800c-b869-3c09f10ad685&#34;&gt;Ref&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Newton&amp;rsquo;s discovery of the universal law of gravitation&lt;/li&gt;
&lt;li&gt;Einstein&amp;rsquo;s discovery of General Relativity&lt;/li&gt;
&lt;li&gt;McClintock&amp;rsquo;s discovery of Transposable Elements: genes that can turn physical characteristics on and off&lt;/li&gt;
&lt;li&gt;Mullis&amp;rsquo; invention of the PCR that makes billions of DNA copies rapidly&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://arxiv.org/abs/2410.12851&#34;&gt;VibeCheck&lt;/a&gt; can predict a model based on its vibes 80% of the time.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.answer.ai/posts/2024-09-03-llmstxt.html&#34;&gt;/llms.txt&lt;/a&gt; is a proposal to standardize &lt;code&gt;/llms.txt&lt;/code&gt; files as a way to share LLM prompts.
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://docs.jina.ai/&#34;&gt;Jina AI Meta Prompt&lt;/a&gt; is an example&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.remotion.dev/docs/system-prompt&#34;&gt;Remotion system prompt&lt;/a&gt; is an example&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://docs.fastht.ml/llms-ctx.txt&#34;&gt;https://docs.fastht.ml/llms-ctx.txt&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://docs.fastht.ml/llms-ctx-full.txt&#34;&gt;https://docs.fastht.ml/llms-ctx-full.txt&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://developer.mozilla.org/en-US/docs/Web/API/Window/structuredClone&#34;&gt;structuredClone&lt;/a&gt; deep clones objects in JS&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/SWivid/F5-TTS&#34;&gt;F5-TTS&lt;/a&gt; clones voices with just 15-second samples.&lt;/li&gt;
&lt;li&gt;Rust has crazy low memory usage too. Spawning thousands of child processes is common and OK these days. &lt;a href=&#34;https://github.com/pretzelhammer/rust-blog/blob/master/posts/rust-in-non-rust-servers.md&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;SetInterval is a good idea in cyborg scraping. &lt;a href=&#34;https://til.simonwillison.net/twitter/collecting-replies&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;GH CLI is quite good for deployment too, like Wrangler CLI. Enabling pages, setting secrets, etc.&lt;/li&gt;
&lt;li&gt;Restic is a CLI backup tool. Just like git. Works well with rclone.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/meta-llama/llama-recipes/tree/main/recipes/quickstart/NotebookLlama&#34;&gt;NotebookLlama&lt;/a&gt; is an open source podcast generator like NotebookLM&lt;/li&gt;
&lt;li&gt;Pragmatic Podcast (I forgot which one)
&lt;ul&gt;
&lt;li&gt;Automate changelogs for your codebases. Convert past commits into attractive release notes automatically&lt;/li&gt;
&lt;li&gt;AI is going to be the consumer of many tools and logs. Build converters for these&lt;/li&gt;
&lt;li&gt;Speed of validation such as linting, testing, etc. will allow LLMs to iterate faster and WILL become more important&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Via Soumya Ranjan
&lt;ul&gt;
&lt;li&gt;Vision embedding is useful in agile modeling&lt;/li&gt;
&lt;li&gt;Vision embedding models with SAM, Grounding Dino by meta, Alibaba does good stuff&lt;/li&gt;
&lt;li&gt;Vision embedding is more useful in batch than real time&lt;/li&gt;
&lt;li&gt;Embedding subtraction with vision embedding models like Dino&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;AI code editors are not good with large code bases today. Keep the refactoring exercises to below 1000 lines. Also evaluate the ease of setting it up locally&lt;/li&gt;
&lt;li&gt;Deepseek Janus is a 1.3b model that can generate both text AND images (and also supports vision)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://cohere.com/blog/multimodal-embed-3&#34;&gt;Cohere Multimodal Embed v3&lt;/a&gt; is available on Azure.&lt;/li&gt;
&lt;li&gt;Elevenlabs lets you create voices with a prompt. No need to even clone one!&lt;/li&gt;
&lt;li&gt;Runway Act One creates expressive character performances&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Clone any voice with a 15-second sample</title>
      <link>https://www.s-anand.net/blog/clone-any-voice-with-a-15-second-sample/</link>
      <pubDate>Thu, 24 Oct 2024 01:36:21 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/clone-any-voice-with-a-15-second-sample/</guid>
      <description>&lt;p&gt;&lt;img alt=&#34;Clone any voice with a 15-second sample&#34; loading=&#34;lazy&#34; src=&#34;https://www.s-anand.net/blog/assets/calvin-voice-cloning.webp&#34;&gt;&lt;/p&gt;
&lt;p&gt;It&#39;s surprisingly easy to clone a voice using &lt;a href=&#34;https://github.com/SWivid/F5-TTS&#34;&gt;F5-TTS&lt;/a&gt;: &#34;A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching&#34;.&lt;/p&gt;
&lt;p&gt;Here&#39;s a clip of me, saying:&lt;/p&gt;
&lt;blockquote class=&#34;wp-block-quote&#34;&gt;
&lt;p&gt;I think Taylor Swift is the best singer. I&#39;ve attended every one of her concerts and in fact, I&#39;ve even proposed to her once. Don&#39;t tell anyone.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(Which is ironic since I didn&#39;t know who she was until this year and I still haven&#39;t seen or heard her.)&lt;/p&gt;
&lt;figure class=&#34;wp-block-audio&#34;&gt;&lt;audio controls=&#34;&#34; src=&#34;https://www.s-anand.net/blog/assets/anand-proposes-to-taylor-swift.opus&#34;&gt;&lt;/audio&gt;&lt;/figure&gt;
&lt;p&gt;You&#39;ll notice that my voice is a bit monotic. That&#39;s because I trained it on a segment of my talk that&#39;s monotonic.&lt;/p&gt;
&lt;figure class=&#34;wp-block-audio&#34;&gt;&lt;audio controls=&#34;&#34; src=&#34;https://www.s-anand.net/blog/assets/anand.opus&#34;&gt;&lt;/audio&gt;&lt;/figure&gt;
&lt;p&gt;&lt;a href=&#34;https://colab.research.google.com/drive/1YEMIdby-Nr--ox2-Wl5Racb5f0ncpqRu?usp=sharing&#34;&gt;&lt;strong&gt;Here&#39;s the code&lt;/strong&gt;&lt;/a&gt;. You can run this on Google Colab for free.&lt;/p&gt;
&lt;p&gt;A few things to keep in mind when preparing the audio.&lt;/p&gt;
&lt;ol class=&#34;wp-block-list&#34;&gt;
&lt;li&gt;Keep the input to just under 15 seconds. That&#39;s the optimal length&lt;/li&gt;
&lt;li&gt;For expressive output, use an input with a broad range of voice emotions&lt;/li&gt;
&lt;li&gt;When using unusual words (e.g. LLM), including the word in your sample helps&lt;/li&gt;
&lt;li&gt;Transcribe &lt;code&gt;input.txt&lt;/code&gt; &lt;em&gt;manually&lt;/em&gt; to get it right, though &lt;a href=&#34;https://huggingface.co/openai/whisper-large-v3&#34;&gt;Whisper&lt;/a&gt; is fine to clone in bulk. (But then, who are you and &lt;em&gt;what&lt;/em&gt; are you doing?)&lt;/li&gt;
&lt;li&gt;Sometimes, each chunk of audio generated has a second of audio from the original interspersed. I don&#39;t know why. Maybe a second of silence at the end helps&lt;/li&gt;
&lt;li&gt;Keep punctuation simple in the generated text. For example, avoid hyphens like &#34;This is obvious - don&#39;t try it.&#34; Use &#34;This is obvious, don&#39;t try it.&#34; instead.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This has a number of uses I can think of (er... ChatGPT can think of), but the ones I find most interesting are:&lt;/p&gt;
&lt;ol class=&#34;wp-block-list&#34;&gt;
&lt;li&gt;&lt;strong&gt;Author-narrated audio books&lt;/strong&gt;. I&#39;m sure this is coming soon, if it&#39;s not already there.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Personalized IVR&lt;/strong&gt;. Why should &lt;em&gt;my&lt;/em&gt; IVR speak in some &lt;em&gt;other&lt;/em&gt; robot&#39;s voice? Let&#39;s use mine. (This has some prank potential.)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Annotated presentations&lt;/strong&gt;. I&#39;m too lazy to speak. Typing is easier. This lets me create, for example, slide decks with my voice, but with &lt;em&gt;editing made super-easy&lt;/em&gt;. I just change the text and the audio changes.&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 03 Mar 2024</title>
      <link>https://www.s-anand.net/blog/things-i-learned-03-mar-2024/</link>
      <pubDate>Sun, 03 Mar 2024 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-03-mar-2024/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://lamplightdev.com/blog/2024/01/10/streaming-html-out-of-order-without-javascript/&#34;&gt;You can use slots to stream HTML out of order&lt;/a&gt;!&lt;/li&gt;
&lt;li&gt;Shane Parrish. Short-term patience podcast
&lt;ul&gt;
&lt;li&gt;have a frame of reference to relate EVERY experience to. That helps you evaluate (measure) and learn. That&amp;rsquo;s part of what Charlie Munger&amp;rsquo;s lattice of frameworks is about&lt;/li&gt;
&lt;li&gt;when there is a very high or very low interest scenario, low interest scenario then go ultra long term. Issued hundred years when the interest rate regime was very low&lt;/li&gt;
&lt;li&gt;short term optimal is rally long term optimal. So you need to learn to take a loss and look like an idiot to play the long-term game&lt;/li&gt;
&lt;li&gt;grit is a behavior that enables long-term thinking. Short term success gives you the luxury to think about long term&lt;/li&gt;
&lt;li&gt;#IMP power is about optionality. It&amp;rsquo;s about being in a position where you have the options that can affect the positive change rather than circumstances controlling you. Read Robert greene&amp;rsquo;s book on the 48 laws of Power&lt;/li&gt;
&lt;li&gt;low leverage enables that&lt;/li&gt;
&lt;li&gt;begin with the end in mind. Always&lt;/li&gt;
&lt;li&gt;how do you think about risk? Well, things do happen. It&amp;rsquo;s as simple as that&lt;/li&gt;
&lt;li&gt;autonomy and decentralization helps derisk&lt;/li&gt;
&lt;li&gt;do more and more of what works. That&amp;rsquo;s a powerful way of compounding&lt;/li&gt;
&lt;li&gt;long-term investments are better than frequent trading because you get to reinvest the tax you otherwise would have paid. So unless the alternative is super compelling, stay invested&lt;/li&gt;
&lt;li&gt;if you need to be the person who DOES the thing, you delegate less, leverage list, compound less, because you have to DO. BE A PERSON WHO SETS THE FIELD INSTEAD. The coach, the chess master, the director, patient strategist who Waits for the good hit&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Being in Control motivates #Lesson. my cycle tires were flat. I thought it was someone pulling out the air and felt very demotivated. But once I carried my cycle pump, I felt so much more in control and power and felt a whole lot better&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://sourcegraph.com/&#34;&gt;SourceGraph&lt;/a&gt; is the default platform for private code completion &amp;amp; search&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.geeky-gadgets.com/ai-voice-cloning-and-creation/&#34;&gt;MetaVoice 1B&lt;/a&gt; offers voice cloning on American &amp;amp; British accents with 30s training&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://qwenlm.github.io/blog/qwen1.5/&#34;&gt;Qwen 1.5 72B&lt;/a&gt; appears to outperform Mistral Medium, making it one of the top non-proprietary models&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://llava-vl.github.io/blog/2024-01-30-llava-1-6/&#34;&gt;Llava 1.6&lt;/a&gt; is a substantial improvement over Llava 1.5 and slightly better than CogVLM, Qwen-VL&lt;/li&gt;
&lt;li&gt;AI scams are growing. &lt;a href=&#34;https://www.straitstimes.com/asia/east-asia/hk-firm-scammed-of-34-million-after-employee-is-duped-by-video-call-with-deepfake-of-cfo&#34;&gt;Deepfakes scammed $34m&lt;/a&gt;. But &lt;a href=&#34;https://www.sfchronicle.com/bayarea/article/ai-phone-scam-18561537.php?sid=64ffe30738148943ca040a9b&amp;amp;ss=A&amp;amp;st_rid=40d8ca22-ad29-44d1-bcfa-45dcd33455f0&#34;&gt;voice fake for kidnapping&lt;/a&gt; is scarier.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://youtu.be/u9dZd4jIxL0&#34;&gt;Buildspace&amp;rsquo;s demo&lt;/a&gt; is a great demo of how voice and actions can be used effectively.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/adefossez/demucs&#34;&gt;demucs&lt;/a&gt; does an EXCELLENT job of splitting songs into drums, bass, vocals and others&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
  </channel>
</rss>
