--- title: Inference description: AI notes by Mrinaal Arora tagged "Inference" updated: 2026-09-24 canonical: https://aroramrinaal.com/ai/tags/inference source: https://aroramrinaal.com/ai/tags/inference.md alternate: https://aroramrinaal.com/ai/tags/inference/markdown --- # Inference AI notes, small builds, and writeups by Mrinaal Arora tagged "Inference". - Page: https://aroramrinaal.com/ai/tags/inference - Markdown: https://aroramrinaal.com/ai/tags/inference/markdown - Markdown (.md): https://aroramrinaal.com/ai/tags/inference.md - All tags: https://aroramrinaal.com/ai/tags/markdown - Full AI index: https://aroramrinaal.com/ai/markdown - Posts with this tag: 3 of 19 - Co-occurring tags: Autoresearch, ICML ## Posts - [Rumik TTS Fast Inference](https://aroramrinaal.com/ai/rumik-tts-fast-inference/markdown) — 2026-09-24; tags: Inference. 60+ experiments to make Rumik OSS 1 faster: 10.74× the stock inference speed on one H100, with automated speech-quality screening. [HTML](https://aroramrinaal.com/ai/rumik-tts-fast-inference) - [Hybrid Token-Efficient Routing Agent](https://aroramrinaal.com/ai/hybrid-token-efficient-routing-agent/markdown) — 2026-07-22; tags: Autoresearch, Inference. How I built TOKENMAN for AMD Developer Hackathon ACT II: a deterministic, local-model, and Fireworks routing agent that climbed from 6,101 counted tokens to zero. [HTML](https://aroramrinaal.com/ai/hybrid-token-efficient-routing-agent) - [\[ICML ’26 Effort\] Efficient Qwen: Making Qwen3.5-4B Faster on a Single A10G](https://aroramrinaal.com/ai/icml-adaptfm-efficient-qwen/markdown) — 2026-06-30; tags: ICML, Autoresearch, Inference. My ICML 2026 AdaptFM effort to optimize Qwen3.5-4B for low-latency inference on a single NVIDIA A10G through 100+ runtime, compiler, and quantization experiments. [HTML](https://aroramrinaal.com/ai/icml-adaptfm-efficient-qwen)