<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>deepseek on S Anand</title>
    <link>https://www.s-anand.net/blog/tag/deepseek/</link>
    <description>Recent content in deepseek on S Anand</description>
    <generator>Hugo -- 0.164.0</generator>
    <language>en-us</language>
    <lastBuildDate>Sun, 26 Jan 2025 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://www.s-anand.net/blog/tag/deepseek/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Things I Learned - 26 Jan 2025</title>
      <link>https://www.s-anand.net/blog/things-i-learned-26-jan-2025/</link>
      <pubDate>Sun, 26 Jan 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-26-jan-2025/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Something I learned from a Sikkil Gurucharan concert.
&lt;ul&gt;
&lt;li&gt;Make the subject of your talk the hero. Not yourself. Be a fan. Share your enthusiasm&lt;/li&gt;
&lt;li&gt;Get into the zone while presenting.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;We reject opposite world views. It&amp;rsquo;s too much effort. But exposure reduces effort and can let us see things from other points of view. So expose yourself to difficult alternative perspectives. &lt;a href=&#34;https://gemini.google.com/share/0a567488cc7a&#34;&gt;Gemini&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Something I learnt from &lt;a href=&#34;https://youtu.be/AjoQTODx0rY&#34;&gt;Aboorva Singeetham&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;Kamal Hassan: &amp;ldquo;A farmer invests in crops. I&amp;rsquo;m an actor. So I invest in films.&amp;rdquo; As a technologist, I guess I would invest in technology.&lt;/li&gt;
&lt;li&gt;&amp;ldquo;A person who has much more to give is unfazed by overwhelming demands because there is too much in him to overwhelm. He gives you 2 options in place of one.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;According to &lt;a href=&#34;https://docsend.com/view/wei3digde8cvmwsr&#34;&gt;Portkey&amp;rsquo;s LLM usage analysis&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Anyscale and Fireworks AI have the lowest error rates (5xx, 429) and rate limits across providers&lt;/li&gt;
&lt;li&gt;Groq and Anthropic are among the highest, OpenAI is among the lowest, Google is in-between&lt;/li&gt;
&lt;li&gt;OpenAI has lower error rates and lower latency than Azure&lt;/li&gt;
&lt;li&gt;They have a ~35% cache hit rate&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A few quick points supporting the mental model of &amp;ldquo;LLMs are aliens&amp;rdquo;.
&lt;ul&gt;
&lt;li&gt;LLMs are clearly not machines. They give different answers each time.&lt;/li&gt;
&lt;li&gt;LLMs &lt;em&gt;are&lt;/em&gt; like humans: they exhibit human biases (e.g. guessing 42 or 37 often). But they fail in unusual ways. They can&amp;rsquo;t count the &amp;ldquo;r&amp;quot;s in strawberry. They can go into an endless loop.&lt;/li&gt;
&lt;li&gt;LLMs are a new form of intelligence. Thinking of them as aliens might minimize our confusions.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Lessons from &lt;a href=&#34;https://www.goodreads.com/book/show/75665850-clear-thinking&#34;&gt;Clear Thinking&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Watch out for four things: Emotion, Ego, Social confirmation, and Inertia/habit. Basically: adrenaline, testosterone, oxytocin, and dopamine. When you feel these, consider doing the opposite.&lt;/li&gt;
&lt;li&gt;Here&amp;rsquo;s what makes us prone to emotion. Sleep deprivation. Hunger. Unknown places. Fatigue. Distraction. Stress (e.g. feeling rushed).&lt;/li&gt;
&lt;li&gt;A good signal for ego is blinding you: You often feel you&amp;rsquo;re right. Or feel unfairly treated.&lt;/li&gt;
&lt;li&gt;Changing behaviors is hard. Instead, join a group or environment where that&amp;rsquo;s the default behavior. Hiring a trainer or joining a gym, for example.&lt;/li&gt;
&lt;li&gt;Why does so much of success literature focus inwards rather than on the environment? Perhaps because we often fool ourselves, and doing less of that gives the biggest bang for the buck. It doesn&amp;rsquo;t mean the environment is unimportant.&lt;/li&gt;
&lt;li&gt;Doing work has the characteristics of a drug. E.g. replying emails gives you control, connections, etc. Work addiction exists because it gives you all the right chemicals.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;If you put LLMs in a feedback loop, it can optimize for its reward function by emotionally pushing people, generating misinformation, nudging towards a narrow definition of creativity, etc.: &lt;a href=&#34;https://bsky.app/profile/emollick.bsky.social/post/3lg4darqwfc2d&#34;&gt;https://bsky.app/profile/emollick.bsky.social/post/3lg4darqwfc2d&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;ChatGPT&amp;rsquo;s &lt;a href=&#34;https://help.openai.com/en/articles/10291617-scheduled-tasks-in-chatgpt&#34;&gt;Scheduled Tasks&lt;/a&gt; are pretty bad at fetching the latest news. Its use of search is poor. (I&amp;rsquo;m not sure if it actually searches.) I need to figure out other use cases for it. Possible options are:&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://api-docs.deepseek.com/quick_start/rate_limit&#34;&gt;DeepSeek does not enforce rate limits&lt;/a&gt;. Yet another reason to switch to DeepSeek. (via &lt;a href=&#34;https://simonwillison.net/2025/Jan/18/deepseek-api-docs-rate-limit/&#34;&gt;Simon Willison&lt;/a&gt;). My other reasons are:
&lt;ul&gt;
&lt;li&gt;Claude 3.5 Sonnet-level coding capability at 5% of the cost (soon to be 2.5%)&lt;/li&gt;
&lt;li&gt;Prompt caching by default&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://api-docs.deepseek.com/guides/fim_completion&#34;&gt;Fill in the middle&lt;/a&gt; completion&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 05 Jan 2025</title>
      <link>https://www.s-anand.net/blog/things-i-learned-05-jan-2025/</link>
      <pubDate>Sun, 05 Jan 2025 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-05-jan-2025/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Some management philosophies used to be successful but are no longer as effective. &lt;a href=&#34;https://chatgpt.com/share/6778b5f9-f1e4-800c-a1f7-30dcdfdccdaa&#34;&gt;ChatGPT&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Command-and-control hierarchy&lt;/li&gt;
&lt;li&gt;Taylorism: deep specialization&lt;/li&gt;
&lt;li&gt;Seniority-based advancement&lt;/li&gt;
&lt;li&gt;Annual performance reviews (without continuous feedback)&lt;/li&gt;
&lt;li&gt;Up-or-Out promotion models&lt;/li&gt;
&lt;li&gt;Confidential strategic information&lt;/li&gt;
&lt;li&gt;Narrow job descriptions&lt;/li&gt;
&lt;li&gt;Relying on formal authority&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Some management philosophies have been around for millenia. &lt;a href=&#34;https://chatgpt.com/share/6778b5f9-f1e4-800c-a1f7-30dcdfdccdaa&#34;&gt;ChatGPT&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Lead by example&lt;/li&gt;
&lt;li&gt;Fairness and empathy&lt;/li&gt;
&lt;li&gt;Clear, consistent communication&lt;/li&gt;
&lt;li&gt;Delegation and empowerment&lt;/li&gt;
&lt;li&gt;Strategic planning and foresight&lt;/li&gt;
&lt;li&gt;Consistent rule enforcement&lt;/li&gt;
&lt;li&gt;Rewarding merit&lt;/li&gt;
&lt;li&gt;Leadership by virtue and character&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.chinatalk.media/p/deepseek-ceo-interview-with-chinas&#34;&gt;Interview with Liang Wenfeng, CEO of DeepSeek&lt;/a&gt;:
&lt;blockquote&gt;
&lt;p&gt;In the face of disruptive technologies, moats created by closed source are temporary. Even OpenAI’s closed source approach can’t prevent others from catching up. So we anchor our value in our team &amp;ndash; our colleagues grow through this process, accumulate know-how, and form an organization and culture capable of innovation. That’s our moat.&lt;/p&gt;
&lt;p&gt;Open source, publishing papers, in fact, do not cost us anything. For technical talent, having others follow your innovation gives a great sense of accomplishment. In fact, open source is more of a cultural behavior than a commercial one, and contributing to it earns us respect. There is also a cultural attraction for a company to do this.&lt;/p&gt;
&lt;p&gt;Why is Silicon Valley so innovative? Because they dare to do things. When ChatGPT came out, the tech community in China lacked confidence in frontier innovation. From investors to big tech, they all thought that the gap was too big and opted to focus on applications instead. But innovation starts with confidence, which we often see more from young people.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://mitmproxy.org/&#34;&gt;mitmproxy&lt;/a&gt; is an open source tool to intercept, modify, and replay HTTP requests. An alternative to &lt;a href=&#34;https://www.charlesproxy.com/&#34;&gt;Charles&lt;/a&gt;, &lt;a href=&#34;https://www.telerik.com/fiddler&#34;&gt;Fiddler&lt;/a&gt;, and partly &lt;a href=&#34;https://www.wireshark.org/&#34;&gt;WireShark&lt;/a&gt;. &lt;a href=&#34;https://earthly.dev/blog/mitmproxy/&#34;&gt;Guide&lt;/a&gt;. Like the others, it requires installing a trusted root certificate on your machine.
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/alufers/mitmproxy2swagger&#34;&gt;mitmproxy2swagger&lt;/a&gt; digs through the mitmproxy flows and generates an OpenAPI schema. A clever idea to reverse-engineer APIs.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://matomo.org/&#34;&gt;Matomo&lt;/a&gt;, &lt;a href=&#34;https://posthog.com/&#34;&gt;PostHog&lt;/a&gt;, &lt;a href=&#34;https://umami.is/&#34;&gt;Umami&lt;/a&gt; and &lt;a href=&#34;https://plausible.io/&#34;&gt;Plausible&lt;/a&gt; are open source web analytics tools (like Google Analytics).&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://redash.io/&#34;&gt;Redash&lt;/a&gt; and &lt;a href=&#34;https://www.metabase.com/&#34;&gt;Metabase&lt;/a&gt; are new open source data visualization tools sitting alongside &lt;a href=&#34;https://grafana.com/&#34;&gt;Grafana&lt;/a&gt; and &lt;a href=&#34;https://superset.apache.org/&#34;&gt;Apache Superset&lt;/a&gt;.
&lt;ul&gt;
&lt;li&gt;Redash feels too clunky / enterprise-y rather than open-source-y.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;From &lt;a href=&#34;https://www.goodreads.com/book/show/27036528-ego-is-the-enemy&#34;&gt;Ego is the enemy&lt;/a&gt;:
&lt;ul&gt;
&lt;li&gt;Add a daily habit to understand your ego. Where and how is it showing up? How are you fooling yourself? Where are you fighting battles without knowing the war?&lt;/li&gt;
&lt;li&gt;Speak less. Do more. E.g. Release more, blog less. Review, THEN publish.&lt;/li&gt;
&lt;li&gt;Always have a teacher, a student, and a peer to compete with. That&amp;rsquo;s how you learn.&lt;/li&gt;
&lt;li&gt;&amp;ldquo;It is impossible for a man to learn what he thinks he already knows&amp;rdquo; - Epictetus.&lt;/li&gt;
&lt;li&gt;Passion makes you blind. Purpose and realism are less so. Delegate, take help, take feedback.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.assemblyai.com/&#34;&gt;Assembly AI&lt;/a&gt; offers speech to text with diarization at 12c/hour. Good diarization, average transcription quality.
In comparison, WhisperX (with GPU) was much slower, had slightly poorer diarization, and slightly better transcription.
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-bash&#34; data-lang=&#34;bash&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;uvx --python 3.9 --index https://download.pytorch.org/whl/cu121 whisperx --diarize --lang en --hf_token &lt;span class=&#34;nv&#34;&gt;$HUGGINGFACE_TOKEN&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://superlinked.com/vector-db-comparison&#34;&gt;Vector DB comparison&lt;/a&gt; compares all popular vector DBs. LanceDB is gently nudging up my preference list but DuckDB is still my favourite.&lt;/li&gt;
&lt;li&gt;Does the cost of &amp;lsquo;running&amp;rsquo; a paper/article in an LLM vary depending on the specific LLM used, such as Claude Sonnet? (FAQ)
&lt;ul&gt;
&lt;li&gt;Yes, the cost varies depending on the LLM. You can see costs and quality at &lt;a href=&#34;https://llmpricing.straive.app/&#34;&gt;https://llmpricing.straive.app/&lt;/a&gt;. The cost is measured in millions of tokens. For example, the Wikipedia page on the Bible is 100K tokens. You can paste text into &lt;a href=&#34;https://platform.openai.com/tokenizer&#34;&gt;https://platform.openai.com/tokenizer&lt;/a&gt; to count the number of tokens.Claude 3.5 Sonnet costs $3.5 / MTok, i.e. 35 cents for the Wikipedia Bible page. Gemini 1.5 Flash 8b costs $0.0375 / MTok, i.e. 0.375 cents for the same page. As you can see, the cost can vary by a factor of 100.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;A git repo with a submodule stores the specific commit of the submodule. When you update the submodule, you need to &lt;code&gt;git add&lt;/code&gt; the submodule.
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;git pull --recurse-submodules&lt;/code&gt;: Pulls parent repo along with submodules&lt;/li&gt;
&lt;li&gt;&lt;code&gt;git submodule status&lt;/code&gt;: For each submodule, show current commit, path, and branch (if on a branch)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;git submodule update --init --recursive&lt;/code&gt;: Fetches/moves each submodule to the commit tracked by the parent repo&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://pypi.org/project/doc2docx/&#34;&gt;&lt;code&gt;uvx doc2docx&lt;/code&gt;&lt;/a&gt; converts Word &lt;code&gt;.doc&lt;/code&gt; files to the new &lt;code&gt;.docx&lt;/code&gt; format. I had several old &lt;code&gt;.doc&lt;/code&gt; files that I converted.&lt;/li&gt;
&lt;li&gt;Sometimes, the value of reading a book is not what you learn from it. It is the thoughts that pop into your head &lt;em&gt;while&lt;/em&gt; reading the book.&lt;/li&gt;
&lt;li&gt;Tools that convert files to prompt / Markdown suitable for LLMs:
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://pypi.org/project/files-to-prompt&#34;&gt;&lt;code&gt;uvx files-to-prompt&lt;/code&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://gitingest.com/&#34;&gt;&lt;code&gt;npx git-ingest&lt;/code&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/sammcj/ingest&#34;&gt;&lt;code&gt;ingest&lt;/code&gt;&lt;/a&gt; - written in Go, only Mac/Linux binaries&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;LLM Code Execution Sandboxes that let you run code in a sandbox via an API:
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/tjmlabs/AgentRun&#34;&gt;AgentRun&lt;/a&gt;: open source, via Docker&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://e2b.dev/&#34;&gt;e2b.dev&lt;/a&gt;: A day costs about $1 (on demand) and you get about $100 one time credits. Self-hosting is complex. &lt;a href=&#34;https://www.reddit.com/r/LocalLLaMA/comments/1chsx7z/is_there_an_opensource_alternative_to_e2b_e2bdev/&#34;&gt;Discussion&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/google/nsjail&#34;&gt;nsjail&lt;/a&gt;: by Google. Write your own API&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;LLM Observability tools:
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://docs.litellm.ai/&#34;&gt;LiteLLM&lt;/a&gt; is an LLM Proxy with caching, logging, call hooks or plugins, rate limiting, virtual keys. SSO integration can be implemented.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://langfuse.com/self-hosting&#34;&gt;LangFuse&lt;/a&gt; is an LLM Proxy with API key distribution, logging, and SSO. But lacks per-user usage limits and server-side caching.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.helicone.ai/&#34;&gt;Helicone&lt;/a&gt; does not support SSO&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://githubnext.com/projects/github-spark&#34;&gt;GitHub Spark&lt;/a&gt; is a way to build micro-apps with LLMs. Like Claude Artifacts. It&amp;rsquo;s currently in technical preview, though.&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    <item>
      <title>Things I Learned - 11 Aug 2024</title>
      <link>https://www.s-anand.net/blog/things-i-learned-11-aug-2024/</link>
      <pubDate>Sun, 11 Aug 2024 00:00:00 +0000</pubDate>
      <guid>https://www.s-anand.net/blog/things-i-learned-11-aug-2024/</guid>
      <description>&lt;p&gt;This week, I learned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Embedding models can be fine-tuned. Example: #TODO&lt;/li&gt;
&lt;li&gt;Agentic RAG (Ravi Theja, LlamaIndex)
&lt;ul&gt;
&lt;li&gt;RAG via top-k retrieval fails with
&lt;ul&gt;
&lt;li&gt;summarization =&amp;gt; need to read all chunks&lt;/li&gt;
&lt;li&gt;comparison: compare product X vs Y =&amp;gt; need to split and re-combine&lt;/li&gt;
&lt;li&gt;structured analytics. e.g. most expensive employees =&amp;gt; Text2SQL first&lt;/li&gt;
&lt;li&gt;multi-part questions. e.g. Tell me about speed of model X AND cost of model Y and recommend =&amp;gt; need to split and re-combine&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;RAG failures: It&amp;rsquo;s single shot. No query planning. No tools. No correction. No memory.&lt;/li&gt;
&lt;li&gt;Agents that help in RAG
&lt;ul&gt;
&lt;li&gt;Route to the right tool
&lt;ul&gt;
&lt;li&gt;E.g. retrieve via vector top-k search or vector summary search or keyword search or combination?&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;One-shot query planning
&lt;ul&gt;
&lt;li&gt;E.g. Break query into multiple specific queries. RAG those. Then combine. #TRY - maybe in DocSearch&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Tool use
&lt;ul&gt;
&lt;li&gt;E.g. Schema retrieval, Text2SQL, Calendar, Chat, APIs, Search, etc.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Agent orchestration
&lt;ul&gt;
&lt;li&gt;ReAct: An agent reasoning loop. Reason + Act. {Thought, Action, Action Input, Observation}*.
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://docs.llamaindex.ai/en/stable/examples/agent/react_agent/&#34;&gt;Orchestrate tools with a prompt&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Multi-agent task solver: &lt;a href=&#34;https://github.com/run-llama/llama-agents&#34;&gt;Llama agents&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Instead of a single agent loop, use different agents. Also allows parallelization&lt;/li&gt;
&lt;li&gt;Allow services to register. (MS TaskWeaver stores tool descriptions in YAML)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://llamahub.ai/?tab=tools&#34;&gt;LlamaHub Tools&lt;/a&gt; has ideas for agents&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes on LLM Fine-Tuning
&lt;ul&gt;
&lt;li&gt;Rouge 2 and Bleu and such metrics are NOT good. Create you own benchmarks&lt;/li&gt;
&lt;li&gt;Non-PEFT fine tuning needs 6X GPU RAM. Optimizer states, Gradient, Activations are the overhead. PEFT is about tuning a subset of parameters.&lt;/li&gt;
&lt;li&gt;LORA adds additional weights without updating the model. It&amp;rsquo;s a low rank matrix multiplication. You can change these adapters in runtime. Saves space. Fast to train&lt;/li&gt;
&lt;li&gt;Quantization: Stick to bitsandbytes or AWQ (may be a bit better)&lt;/li&gt;
&lt;li&gt;QLORA = Quantization + LORA&lt;/li&gt;
&lt;li&gt;Predibase has open-sourced Lora Adapters in &amp;ldquo;Lora Land&amp;rdquo;. Existing adapters are pretty good.
&lt;ul&gt;
&lt;li&gt;ghcr.io/predibase/lorax:main Docker image works on Docker compose to run locally. &lt;code&gt;devices:&lt;/code&gt; on Docker Compose lets you specify NVIDIA GPU devices&lt;/li&gt;
&lt;li&gt;Locust is a HTTP load testing lib in Python&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Techniques for inference optimization
&lt;ul&gt;
&lt;li&gt;Dynamic adapters: Loads right LORAX adapters WHEN a request comes in&lt;/li&gt;
&lt;li&gt;Multi-adapter batching: Process all inputs in parallel on the same GPU, but different users are post-processed using different adapters&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Notes from a 4-hour flight:
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://applied-llms.org/&#34;&gt;What We’ve Learned From A Year of Building with LLMs&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;Strategy
&lt;ul&gt;
&lt;li&gt;IS IT TOO HARD/EXPENSIVE? Log it. LLMs are getting cheaper and better.&lt;/li&gt;
&lt;li&gt;WILL OPENAI BUILD IT? If so, wait for it instead of building.&lt;/li&gt;
&lt;li&gt;HAS A STARTUP BUILT IT? If so, use it instead. It&amp;rsquo;s a generic use case there&amp;rsquo;s no point re-inventing.&lt;/li&gt;
&lt;li&gt;FOCUSED USE CASES over generic. Build trust by starting small.&lt;/li&gt;
&lt;li&gt;Tools for LLM Ops (feedback): LangSmith, Log10, LangFuse, W&amp;amp;B Weave, HoneyHive #TRY&lt;/li&gt;
&lt;li&gt;Human in the Loop is about humans evaluating model outputs. That&amp;rsquo;s different from AI in the loop, human in the center, where AI accelerates human output (like Github Copilot)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Operations
&lt;ul&gt;
&lt;li&gt;CHECK EMBEDDINGS DRIFT over time. Users might be input-ing different things than before.&lt;/li&gt;
&lt;li&gt;LOG AND REVIEW everything.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/jxnl/instructor&#34;&gt;Instructor&lt;/a&gt; coaxes structured output from LLM APIs. #TRY&lt;/li&gt;
&lt;li&gt;IMPLICIT FEEDBACK collection is easy. Just let users edit stuff. #TRY&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Tactical
&lt;ul&gt;
&lt;li&gt;Try n-shot prompting (n=5-12) before bigger models. #TRY&lt;/li&gt;
&lt;li&gt;Always structure for output: Markdown, XML/HTML tags.&lt;/li&gt;
&lt;li&gt;Combine RAG with Keyword search. It reduces user frustration in edge cases.&lt;/li&gt;
&lt;li&gt;Prefer multiple small prompts to one big prompt. Do X. Then Y. Then Z.&lt;/li&gt;
&lt;li&gt;Jitter prompts for diversity beyond temperature.&lt;/li&gt;
&lt;li&gt;LLM-as-judge works better when comparing outputs (not rating 1 output). Keep length similar (LLMs prefer wordiness). Swap order and compare. Allow for ties. Ask for reason FIRST.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://bytes.swiggy.com/hermes-a-text-to-sql-solution-at-swiggy-81573fb4fb6e&#34;&gt;Hermes: A Text-to-SQL solution at Swiggy&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;Hermes performed significantly better for charters with well-defined metadata and a relatively smaller number of tables.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&amp;ldquo;We collect feedback on the accuracy of the returned query from stakeholders directly within the Slack bot.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://nicholas.carlini.com/writing/2024/how-i-use-ai.html&#34;&gt;How I use AI&lt;/a&gt; and &amp;ldquo;Replacing my right hand with AI&amp;rdquo;
&lt;ul&gt;
&lt;li&gt;EMBED in every app/workflow. E.g. Auto-fix spellings. Auto-review code. Auto-ask LLM on errors and apply patch! Auto-search for answer, assess, continue.&lt;/li&gt;
&lt;li&gt;PERSIST. Stick with the LLM to the end. Don&amp;rsquo;t fix it yourself. It&amp;rsquo;s faster. #TRY
&lt;ul&gt;
&lt;li&gt;INTERVENE FAST. If an LLM can&amp;rsquo;t solve it by itself in 2 tries, it needs in-depth help.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;APP-IFY one-off tasks. Disposable tools. &amp;ldquo;Write web-app to convert JSON to tab-delimited.&amp;rdquo; &amp;ldquo;Extract fields as a table.&amp;rdquo; &amp;ldquo;Diff JSON.&amp;rdquo; #TRY&lt;/li&gt;
&lt;li&gt;BEST language/frameworks preferred. CUDA in Python. Rust. C. Raspberry Pi. Arduino. Bluetooth. Modern ESM/JS. #TRY&lt;/li&gt;
&lt;li&gt;TEACH examples. &amp;ldquo;Here&amp;rsquo;s the LLM Foundry API.&amp;rdquo; &amp;ldquo;Here&amp;rsquo;s how to use gramex.data.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;DUMP entire code. Models can handle it. Refactoring to SQLAlchemy 2, Pandas 2. API Documentation. Test case generation. #TRY&lt;/li&gt;
&lt;li&gt;ASK for features &amp;amp; packages. Docker without root access. GPU access inside docker. Windows CLI-only C++ compiler.&lt;/li&gt;
&lt;li&gt;TEST CASE writing. #TRY&lt;/li&gt;
&lt;li&gt;SPEC IN DETAIL. Use these libraries. Write like this: code example.&lt;/li&gt;
&lt;li&gt;SPEC &lt;em&gt;USAGE&lt;/em&gt; in detail.
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;I will just pipe it into sqlite&amp;rdquo;, or &amp;ldquo;I will just run &lt;code&gt;ffmpeg -i filename [YOUR OPTIONS]&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Describe the UI, API input/output, data structure, and internal data structure.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;HELP on usage. &amp;ldquo;ffmpeg to get audio.mp3&amp;rdquo;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://nicholas.carlini.com/writing/2024/my-benchmark-for-large-language-models.html&#34;&gt;My benchmark for large language models&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;LLM(text) is a useful function to have in JS and Python too. Useful as a simple &lt;code&gt;pip install llmfoundry&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Allow images, files in LLM()&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Current list of #IMPOSSIBLE (or hard) things for LLMs
&lt;ul&gt;
&lt;li&gt;Translate technical documents to Dutch &amp;ndash; because they don&amp;rsquo;t understand the technical terms well&lt;/li&gt;
&lt;li&gt;Translate large documents (JSON to XML, English to Chinese, Python to Rust, Wrong to right spelling) &amp;ndash; because the output tokens are limited&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/BuilderIO/micro-agent&#34;&gt;micro-agent&lt;/a&gt; generates test cases first when asked to build an app. Then it iterates until the test cases pass.&lt;/li&gt;
&lt;li&gt;Alternative interfaces to YouTube: Piped.video, CloudTube, Invidious, NewPipe, FreeTube&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://platform.deepseek.com/api-docs/news/news0802/&#34;&gt;Deepseek Context Caching&lt;/a&gt; reduces price to 1.4 cents/MTok for portions of chat messages that are repeated. That&amp;rsquo;s a 10X reduction for long conversations!&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
  </channel>
</rss>
