<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Llm-Inference on socratic notes</title><link>https://socraticagent.dev/tags/llm-inference/</link><description>Recent content in Llm-Inference on socratic notes</description><generator>Hugo</generator><language>en</language><lastBuildDate>Mon, 07 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://socraticagent.dev/tags/llm-inference/index.xml" rel="self" type="application/rss+xml"/><item><title>Why Your LLM Is Slow for Three Completely Different Reasons</title><link>https://socraticagent.dev/blog/why-your-llm-is-slow-for-three-completely-different-reasons/</link><pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate><guid>https://socraticagent.dev/blog/why-your-llm-is-slow-for-three-completely-different-reasons/</guid><description>Time-to-first-token and tokens-per-second aren&amp;#39;t the same problem wearing different names - they&amp;#39;re governed by opposite physical bottlenecks, plus a third one nobody names until requests start queueing for no obvious reason.</description></item></channel></rss>