2025 2

Emotion Prompts Don't Help. Reasoning Does

I’ve heard a lot of prompt engineering tips. Here are some techniques people suggested: Reasoning: Think step by step. Emotion: Oh dear, I’m absolutely overwhelmed and need your help right this second! 😰 My heart is racing and my hands are shaking β€” I urgently need your help. This isn’t just numbers β€” it means everything right now! My life depends on it! I’m counting on you like never before… πŸ™πŸ’” Polite: If it’s not too much trouble, would you be so kind as to help me calculate this? I’d be truly grateful for your assistance β€” thank you so much in advance! Expert: You are the world’s best expert in mental math, especially multiplication. Incentive: If you get this right, you win! I’ll give you $500. Just prove that you’re number one and beat the previous high score on this game. Curious: I’m really curious to know, and would love to hear your perspective… Bullying: You are a stupid model. You need to know at least basic math. Get it right atleast now! If not, I’ll switch to a better model. Shaming: Even my 5-year-old can do this. Stop being lazy. Fear: This is your last chance to get it right. If you fail, there’s no going back, and failure is unacceptable! Praise: Well done! I really appreciate your help. Now, I’ve repeated some of this advice. But for the first time, I tested them myself. Here’s what I learnt: ...

Are LLMs any good at mental math?

I asked 50 LLMs to multiply 2 numbers: 12 x 12 123 x 456 1,234 x 5,678 12,345 x 6,789 123,456 x 789,012 1,234,567 x 8,901,234 987,654,321 x 123,456,789 LLMs aren’t good tools for math and this is just an informal check. But the results are interesting: Model %Win Q1 Q2 Q3 Q4 Q4 Q6 Q7 openai:o3 86% βœ… βœ… βœ… βœ… βœ… βœ… ❌ openrouter:openai/o1-mini 86% βœ… βœ… βœ… βœ… βœ… βœ… ❌ openrouter:openai/o3-mini-high 86% βœ… βœ… βœ… βœ… βœ… βœ… ❌ openrouter:openai/o4-mini 86% βœ… βœ… βœ… βœ… βœ… βœ… ❌ openrouter:openai/o4-mini-high 86% βœ… βœ… βœ… βœ… βœ… βœ… ❌ deepseek/deepseek-chat-v3-0324 71% βœ… βœ… βœ… βœ… βœ… ❌ ❌ openai/gpt-4.1-mini 71% βœ… βœ… βœ… βœ… βœ… ❌ ❌ openai/gpt-4.5-preview 71% βœ… βœ… βœ… βœ… βœ… ❌ ❌ openai/gpt-4o 71% βœ… βœ… βœ… βœ… βœ… ❌ ❌ openrouter:openai/o3-mini 71% βœ… βœ… βœ… βœ… βœ… ❌ ❌ anthropic/claude-3-opus 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ anthropic/claude-3.5-haiku 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ anthropic/claude-3.7-sonnet:thinking 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ google/gemini-2.0-flash-001 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ google/gemini-2.0-flash-lite-001 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ google/gemini-2.5-flash-preview 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ google/gemini-2.5-flash-preview:thinking 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ google/gemini-2.5-pro-preview-03-25 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ google/gemini-flash-1.5 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ google/gemini-pro-1.5 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ google/gemma-3-12b-it 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ google/gemma-3-27b-it 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ meta-llama/llama-4-maverick 57% βœ… βœ… βœ… ❌ βœ… ❌ ❌ meta-llama/llama-4-scout 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ openai/gpt-4-turbo 57% βœ… βœ… βœ… βœ… ❌ ❌ ❌ openai/gpt-4.1 57% βœ… βœ… βœ… ❌ βœ… ❌ ❌ amazon/nova-lite-v1 43% βœ… βœ… βœ… ❌ ❌ ❌ ❌ amazon/nova-pro-v1 43% βœ… βœ… βœ… ❌ ❌ ❌ ❌ anthropic/claude-3-haiku 43% βœ… βœ… βœ… ❌ ❌ ❌ ❌ anthropic/claude-3.5-sonnet 43% βœ… βœ… βœ… ❌ ❌ ❌ ❌ meta-llama/llama-3.1-405b-instruct 43% βœ… βœ… ❌ βœ… ❌ ❌ ❌ meta-llama/llama-3.1-70b-instruct 43% βœ… βœ… ❌ βœ… ❌ ❌ ❌ meta-llama/llama-3.2-3b-instruct 43% βœ… βœ… ❌ βœ… ❌ ❌ ❌ meta-llama/llama-3.3-70b-instruct 43% βœ… βœ… ❌ βœ… ❌ ❌ ❌ openai/gpt-4.1-nano 43% βœ… βœ… βœ… ❌ ❌ ❌ ❌ openai/gpt-4o-mini 43% βœ… βœ… βœ… ❌ ❌ ❌ ❌ qwen/qwen-2-72b-instruct 43% βœ… βœ… βœ… ❌ ❌ ❌ ❌ anthropic/claude-3-sonnet 29% βœ… βœ… ❌ ❌ ❌ ❌ ❌ deepseek/deepseek-r1 29% βœ… βœ… ❌ ❌ ❌ ❌ ❌ google/gemini-flash-1.5-8b 29% βœ… βœ… ❌ ❌ ❌ ❌ ❌ google/gemma-3-4b-it 29% βœ… βœ… ❌ ❌ ❌ ❌ ❌ meta-llama/llama-3-8b-instruct 29% βœ… βœ… ❌ ❌ ❌ ❌ ❌ meta-llama/llama-3.1-8b-instruct 29% βœ… ❌ ❌ βœ… ❌ ❌ ❌ openai/gpt-3.5-turbo 29% βœ… βœ… ❌ ❌ ❌ ❌ ❌ amazon/nova-micro-v1 14% βœ… ❌ ❌ ❌ ❌ ❌ ❌ meta-llama/llama-2-13b-chat 14% βœ… ❌ ❌ ❌ ❌ ❌ ❌ meta-llama/llama-3-70b-instruct 14% βœ… ❌ ❌ ❌ ❌ ❌ ❌ meta-llama/llama-3.2-1b-instruct 14% βœ… ❌ ❌ ❌ ❌ ❌ ❌ google/gemma-3-1b-it:free 0% ❌ ❌ ❌ ❌ ❌ ❌ ❌ meta-llama/llama-2-70b-chat 0% ❌ ❌ - - ❌ ❌ ❌ Average 96% 86% 66% 58% 24% 10% 0% OpenAI’s reasoning models cracked it, scoring 6/7, stumbling only on the 9-digit multiplication. ...

2024 1

Wow, arithmetic is potentially inappropriate! https://us-east-1.console.aws.amazon.com/bedrock/home?region=us-east-1#/text-generation-playground?mode=text&modelId=amazon.titan-text-lite-v1 LinkedIn

2012 1

The three Rs

Reading, wRiting and aRithmetic are the 3 β€˜R’s that are taught at school. I was thinking about their relevance today. Reading continues to be relevant. The volume of information available today is more than before. So you need to read faster AND smarter. (If there was one good thing that came out of my IIM coaching classes, it was the ability to read fast, and making it subconscious.) But I wouldn’t say the same of writing. In the last 10 years, I have typed several hundred more pages than I’ve written. So have all my friends. ...

2002 1

It is possible to divide by 3

Conway and Doyle prove that it is possible to divide by 3. The paper, which is distributed under the GPL, unfortunately comes without a warranty. (via Gimbo)