Model: https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b
- Architecture: Cache-Aware FastConformer encoder (24 layers) + RNNT decoder, 600M params — streaming with configurable 80ms–1120ms chunk latency
- Formats: safetensors, PyTorch, GGUF (q8_0) via NeMo-Speech.cpp (C++ runtime)
- Official hardware/OS support: NVIDIA GPU archs only (Ampere/Blackwell/Hopper/Jetson/Lovelace/Turing/Volta), Linux / Linux4Tegra. No ONNX, Core ML, or TFLite export path officially provided. No mobile or web runtime published.
- 40 language-locales, license OpenMDW-1.1