---
title: Inference
description: AI notes by Mrinaal Arora tagged "Inference"
updated: 2026-09-24
canonical: https://aroramrinaal.com/ai/tags/inference
source: https://aroramrinaal.com/ai/tags/inference.md
alternate: https://aroramrinaal.com/ai/tags/inference/markdown
---

# Inference

AI notes, small builds, and writeups by Mrinaal Arora tagged "Inference".

- Page: https://aroramrinaal.com/ai/tags/inference
- Markdown: https://aroramrinaal.com/ai/tags/inference/markdown
- Markdown (.md): https://aroramrinaal.com/ai/tags/inference.md
- All tags: https://aroramrinaal.com/ai/tags/markdown
- Full AI index: https://aroramrinaal.com/ai/markdown
- Posts with this tag: 3 of 19
- Co-occurring tags: Autoresearch, ICML

## Posts

- [Rumik TTS Fast Inference](https://aroramrinaal.com/ai/rumik-tts-fast-inference/markdown) — 2026-09-24; tags: Inference. 60+ experiments to make Rumik OSS 1 faster: 10.74× the stock inference speed on one H100, with automated speech-quality screening. [HTML](https://aroramrinaal.com/ai/rumik-tts-fast-inference)
- [Hybrid Token-Efficient Routing Agent](https://aroramrinaal.com/ai/hybrid-token-efficient-routing-agent/markdown) — 2026-07-22; tags: Autoresearch, Inference. How I built TOKENMAN for AMD Developer Hackathon ACT II: a deterministic, local-model, and Fireworks routing agent that climbed from 6,101 counted tokens to zero. [HTML](https://aroramrinaal.com/ai/hybrid-token-efficient-routing-agent)
- [\[ICML ’26 Effort\] Efficient Qwen: Making Qwen3.5-4B Faster on a Single A10G](https://aroramrinaal.com/ai/icml-adaptfm-efficient-qwen/markdown) — 2026-06-30; tags: ICML, Autoresearch, Inference. My ICML 2026 AdaptFM effort to optimize Qwen3.5-4B for low-latency inference on a single NVIDIA A10G through 100+ runtime, compiler, and quantization experiments. [HTML](https://aroramrinaal.com/ai/icml-adaptfm-efficient-qwen)
