<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>elevenlabs on S Anand</title>
    <link>https://www.s-anand.net/blog/tag/elevenlabs/</link>
    <description>Recent content in elevenlabs on S Anand</description>
    <generator>Hugo -- 0.164.0</generator>
    <language>en-us</language>
    <lastBuildDate>Sun, 29 Mar 2026 22:31:33 +0800</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/tag/elevenlabs/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>MGR via ElevenLabs</title>
      <link>https://www.s-anand.net/blog/mgr-via-elevenlabs/</link>
      <pubDate>Sun, 29 Mar 2026 22:31:33 +0800</pubDate>
      <guid>https://www.s-anand.net/blog/mgr-via-elevenlabs/</guid>
      <description>&lt;p&gt;I was watching &lt;a href=&#34;https://en.wikipedia.org/wiki/Vaa_Vaathiyaar&#34;&gt;Vaa Vaathiyar&lt;/a&gt; which has a short clip of &lt;a href=&#34;https://en.wikipedia.org/wiki/M._G._Ramachandran&#34;&gt;MGR&lt;/a&gt; speaking. It&amp;rsquo;s either AI-generated or mimic-ed and it wasn&amp;rsquo;t bad.&lt;/p&gt;
&lt;p&gt;&lt;audio controls src=&#34;https://files.s-anand.net/images/2026-03-29-mgr-in-film.opus&#34;&gt;&lt;/audio&gt;&lt;/p&gt;
&lt;p&gt;I used &lt;code&gt;ffmpeg&lt;/code&gt; to record the audio from the film, transcribed it via &lt;a href=&#34;https://aistudio.google.com/prompts/new_chat?model=gemini-3.1-pro-preview&#34;&gt;Gemini 3 Pro on AI Studio&lt;/a&gt; with the prompt:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Transcribe this into Tamil&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&amp;hellip; which gave me:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;ராமு&amp;hellip;
என்ன செய்திருக்கிறாய் நீ&amp;hellip;
வாத்தியார் கேட்கிறேன் சொல்
நிமிர்ந்து பார்க்க கூட தைரியம் இல்லையா&amp;hellip;
ஓடாதே&amp;hellip; நில்&amp;hellip;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Translation:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Ramu&amp;hellip;
What have you done&amp;hellip;
Vaathiyar (MGR) is asking, tell me
Don&amp;rsquo;t you have the courage to stand up and look at me&amp;hellip;
Don&amp;rsquo;t run&amp;hellip; stop&amp;hellip;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(GitHub Copilot&amp;rsquo;s auto-complete translated the above for me as I typed - flawlessly. It&amp;rsquo;s getting better by the day!)&lt;/p&gt;
&lt;p&gt;Then, I used &lt;a href=&#34;https://github.com/yt-dlp/yt-dlp&#34;&gt;yt-dlp&lt;/a&gt; to download the audio from this &lt;a href=&#34;https://www.youtube.com/shorts/1jQqKds2z7g&#34;&gt;MGR Short Clip&lt;/a&gt;.&lt;/p&gt;
&lt;iframe width=&#34;560&#34; height=&#34;315&#34; src=&#34;https://www.youtube.com/embed/1jQqKds2z7g?si=8GUnT7IvcsQZpwDb&#34; title=&#34;YouTube video player&#34; frameborder=&#34;0&#34; allow=&#34;accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share&#34; referrerpolicy=&#34;strict-origin-when-cross-origin&#34; allowfullscreen&gt;&lt;/iframe&gt;
&lt;p&gt;Here&amp;rsquo;s the sample:&lt;/p&gt;
&lt;p&gt;&lt;audio controls src=&#34;https://files.s-anand.net/images/2026-03-29-mgr-sample.opus&#34;&gt;&lt;/audio&gt;&lt;/p&gt;
&lt;p&gt;I fed this into ElevenLabs&amp;rsquo; &lt;a href=&#34;https://elevenlabs.io/app/voice-library&#34;&gt;Instant Voice Clone&lt;/a&gt; that needs just 10 seconds of audio and created an &amp;ldquo;MGR&amp;rdquo; voice.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the same dialogue in the cloned voice:&lt;/p&gt;
&lt;p&gt;&lt;audio controls src=&#34;https://files.s-anand.net/images/2026-03-29-mgr-generated.opus&#34;&gt;&lt;/audio&gt;&lt;/p&gt;
&lt;p&gt;Personally, I think the ElevenLabs version is &lt;em&gt;slightly&lt;/em&gt; better. Of course, given the pace of AI improvement, this might just be the impact of a new model release.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;2 Apr 2026&lt;/strong&gt;: Here&amp;rsquo;s the non-cloned generation from &lt;a href=&#34;https://dashboard.sarvam.ai/text-to-speech&#34;&gt;Sarvam&amp;rsquo;s text to speech&lt;/a&gt; with Bulbul v3 standard quality. It feels pretty weak.&lt;/p&gt;
&lt;p&gt;&lt;audio controls src=&#34;https://files.s-anand.net/images/2026-03-29-mgr-sarvam.opus&#34;&gt;&lt;/audio&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://aistudio.google.com/u/2/generate-speech?model=gemini-2.5-pro-preview-tts&#34;&gt;Gemini 2.5 Pro Preview TTS&lt;/a&gt; gave me this, which feels &lt;em&gt;much&lt;/em&gt; better.&lt;/p&gt;
&lt;p&gt;&lt;audio controls src=&#34;https://files.s-anand.net/images/2026-03-29-mgr-gemini.opus&#34;&gt;&lt;/audio&gt;&lt;/p&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 14 Jul 2024</title>
      <link>https://www.s-anand.net/blog/things-i-learned-14-jul-2024/</link>
      <pubDate>Sun, 14 Jul 2024 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-14-jul-2024/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Carlton&amp;rsquo;s TDS session
&lt;ul&gt;
&lt;li&gt;Always create a new venv via VS Code when starting a training session. Helps reproduce issues (though I could use Colab instead)&lt;/li&gt;
&lt;li&gt;Create an empty .ipynb notebook and double-click it. That&amp;rsquo;s another way (though slower) to open a Jupyter notebook&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Share Parrish Knowledge Project podcast. Three generations of wealth
&lt;ul&gt;
&lt;li&gt;There is a big difference between liking animals and being a vet. Between liking education and being a teacher.&lt;/li&gt;
&lt;li&gt;Even if no one reads your writing, you benefit from the writing.&lt;/li&gt;
&lt;li&gt;Emotional.crises like 9/11 or Covid are far easier for markets to recover from&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Hidden brain podcast. White trying to hard can back fire on you
&lt;ul&gt;
&lt;li&gt;Sometimes conscious thinking makes our automated responses of sports music, dance are great examples&lt;/li&gt;
&lt;li&gt;Instead, SURRENDER to something outside of you. Like playing with kids. Exercise also sends blood away from brain. Drugs. ChatGPT.&lt;/li&gt;
&lt;li&gt;It&amp;rsquo;s called Ue in Chinese philosophy&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A quick check on the pricing of text to speech models
&lt;ul&gt;
&lt;li&gt;OpenAI TTS: $15/1M chars &lt;a href=&#34;https://openai.com/api/pricing/&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Deepgram Aura: $15/1M chars &lt;a href=&#34;https://deepgram.com/pricing&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Elevenlabs Scale: $165/1M chars &lt;a href=&#34;https://elevenlabs.io/pricing&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Google TTS Neural2: $16/1M chars &lt;a href=&#34;https://cloud.google.com/text-to-speech/pricing?hl=en&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Azure AI Speech: $15/1M chars &lt;a href=&#34;https://azure.microsoft.com/en-us/pricing/details/cognitive-services/speech-services/&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;AWS Polly Neural TTS: $16/1M chars &lt;a href=&#34;https://aws.amazon.com/polly/pricing/&#34;&gt;Ref&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 21 Jan 2024</title>
      <link>https://www.s-anand.net/blog/things-i-learned-21-jan-2024/</link>
      <pubDate>Sun, 21 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-21-jan-2024/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;When comparing Mistral with 4b quantization vs unquantized:
&lt;ul&gt;
&lt;li&gt;2 responses were significantly shorter and fairly different&lt;/li&gt;
&lt;li&gt;1 was identical&lt;/li&gt;
&lt;li&gt;1 was almost identical but shorter by a few words&lt;/li&gt;
&lt;li&gt;1 was slightly longer and fairly different&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;#PREDICTION As humans have more conversations with LLMs, they will replace video watching and interactive gaming with conversation based role play. New game genres will evolve&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.lilacml.com/&#34;&gt;Lilac&lt;/a&gt; is an LLM-based data curation tool. Use it to search by concept (e.g. PII, duplicates, etc.) and then drop/update the results.&lt;/li&gt;
&lt;li&gt;Lungs have a Hausdorff dimension of 2.97 &amp;ndash; giving them one of the highest surface area to volume ratio. Brains are 2.8. Sierpinski Pyramid is exactly 2 &amp;ndash; which is weird. To solid-paint twice the size, you need 4 times as much paint.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://youtu.be/rXUuStdMeoE&#34;&gt;How I write podcast. Tim Ferriss&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;High bars are constraints. I set the strongest constraints against the scarcest resources. Like reputation&lt;/li&gt;
&lt;li&gt;Being a category of one is more defensible than a competitive advantage&lt;/li&gt;
&lt;li&gt;Content always beats presentation. When in doubt, push for more interesting content&lt;/li&gt;
&lt;li&gt;Regular publishing improves thinking&lt;/li&gt;
&lt;li&gt;To build a habit, do less than you think you can do. That makes it easier to build momentum on the habit and sustain during crunch times&lt;/li&gt;
&lt;li&gt;There is a lot of mediocrity in the world. If you&amp;rsquo;re doing something (in a winner take all ecosystem), be the best.&lt;/li&gt;
&lt;li&gt;Top lawyers are exceptional proofreaders. They are able to see what is unclair, and what is redundant, and what has loop holes very quickly.&lt;/li&gt;
&lt;li&gt;Forcing yourself to cut down from a thousand words to 200 to a paragraph to a sentence takes you through a phase transition where you discover something unexpected&lt;/li&gt;
&lt;li&gt;The more outrageous the question, the more likely it is to be useful in generating a new perspective&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://elevenlabs.io/speech-synthesis&#34;&gt;Eleven-labs speech synthesis&lt;/a&gt; with voice cloning is at the uncanny valley. With two 5-minute samples, my voice sounds a fair bit like my voice but is very clearly not my voice. I find stability ~ 30%, similarity ~ 80% and style ~50% gives a reasonable outcome. But the default voices (e.g. Joseph, George, Charlie) are excellent.&lt;/li&gt;
&lt;li&gt;Practical AI podcast: AI predictions for&lt;/li&gt;
&lt;li&gt;AI by API is the norm today and will grow
&lt;ul&gt;
&lt;li&gt;Just having AI is no longer a differentiator&lt;/li&gt;
&lt;li&gt;AI is part of life, not just work&lt;/li&gt;
&lt;li&gt;#TODO Explore quickdrop from Stability for Maruti&lt;/li&gt;
&lt;li&gt;#TODO Explore Codium VS Code plugin and Continue.dev&lt;/li&gt;
&lt;li&gt;Hybrid systems that combine stats, ML, DL and AI models will grow&lt;/li&gt;
&lt;li&gt;AGI and AutoGPT resurgence&lt;/li&gt;
&lt;li&gt;RAG will continue to be a focus&lt;/li&gt;
&lt;li&gt;GPT4 will be beaten by open source models. Special purpose models beat it already&lt;/li&gt;
&lt;li&gt;Self hosted and cloud hosted models will grow for security&lt;/li&gt;
&lt;li&gt;Small language models will grow&lt;/li&gt;
&lt;li&gt;Productivity will be enhanced rather than replaced&lt;/li&gt;
&lt;li&gt;Multi modal models will grow&lt;/li&gt;
&lt;li&gt;Cost efficiency will grow in focus&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://help.openai.com/en/articles/8770868-gpt-builder&#34;&gt;GPT Builder help&lt;/a&gt; explains how the GPT Builder updates GPTs - including some very interesting prompts&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
  </channel>
</rss>
