Fast text
to speech.
Rumik OSS 1 text-to-speech inference on an NVIDIA H100.
INT8 weight-only decode, fused Triton kernels, and GPU-controlled CUDA graphs.
The first request can take a few minutes while the GPU starts.
References
- Original model: rumik-ai/rumik-oss-1 on Hugging Face.
- Rumik research: Introducing Rumik OSS 1.
- Base language model: CohereLabs/tiny-aya-fire.
- Audio codec: kyutai/mimi.