I learned how Predibase uses LoRAX to serve multiple fine-tuned model adapters on a single GPU. This method leverages LoRA adapters to optimize hardware usage and scale specialized models more efficiently than traditional deployments.
I explored Agentic RAG for complex retrieval, fine-tuning with LoRAX, and practical LLM strategies. Key takeaways include using N-shot prompting before scaling models, automating workflows via disposable apps, and leveraging context caching to significantly reduce inference costs.
I explored home networking and LLM infrastructure, discovering WiFi 6 beam-forming, Predibase's competitive pricing for fine-tuned models, RunPod's serverless vLLM endpoints for HuggingFace models, and Portkey's utility as an AI model router.
I tested LLMs using Caesar cipher prompts, compiled a list of cheap cloud GPU services like Runpod, and learned how JSR handles package documentation. I also found that averaging embeddings is useful for processing long document inputs.
I explored GPT functions for spreadsheets, Xata's free PostgreSQL API, and Nginx's least_conn load balancing. I also looked into GitHub Copilot's prompt construction and the importance of tracking evolving LLM capabilities and hardware-specific package managers.
I explored building minimal Docker images from scratch, fine-tuned Mistral using Axolotl and Deepspeed, and studied communication strategies for winning hearts. I also practiced D3.js data visualization techniques and integrated Bard with my Google Workspace tools.