back to ai
Hugging Face
home/ai/tags/inference

Inference

Showing 3 of 19 posts tagged Inference

All19AI SafetyAlignmentAutoresearchContinual LearningICMLImage GenerationInferenceInterpretabilityMulti-Agent SystemsMultilingualPre-TrainingRLSFTSmall ModelsSparse Autoencoders

Rumik TTS Fast Inference

60+ experiments to make Rumik OSS 1 faster: 10.74× the stock inference speed on one H100, with automated speech-quality screening.

September 24, 2026

Hybrid Token-Efficient Routing Agent

How I built TOKENMAN for AMD Developer Hackathon ACT II: a deterministic, local-model, and Fireworks routing agent that climbed…

July 22, 2026

[ICML ’26 Effort] Efficient Qwen: Making Qwen3.5-4B Faster on a Single A10G

My ICML 2026 AdaptFM effort to optimize Qwen3.5-4B for low-latency inference on a single NVIDIA A10G through 100+ runtime…

June 30, 2026
September 24, 2026reads

Rumik TTS Fast Inference

60+ experiments to make Rumik OSS 1 faster: 10.74× the stock inference speed on one H100, with automated speech-quality screening.

July 22, 2026reads

Hybrid Token-Efficient Routing Agent

How I built TOKENMAN for AMD Developer Hackathon ACT II: a deterministic, local-model, and Fireworks routing agent that climbed from 6,101 counted tokens to zero.

June 30, 2026reads

[ICML ’26 Effort] Efficient Qwen: Making Qwen3.5-4B Faster on a Single A10G

My ICML 2026 AdaptFM effort to optimize Qwen3.5-4B for low-latency inference on a single NVIDIA A10G through 100+ runtime, compiler, and quantization experiments.