Run optimized text-to-text models from the Arm AI Portal on Arm-powered Android devices
Run optimized text-generation and embedding models from the Arm AI Portal using supported Android runtimes and model adapters.
Run optimized text-generation and embedding models from the Arm AI Portal using supported Android runtimes and model adapters.
Build an end-to-end, on-device voice assistant that understands both speech and emotion using Whisper, HuBERT, ONNX Runtime, and a local LLM with llama.cpp on Arm.
Build and benchmark a multimodal Android voice-assistant pipeline, then use KleidiAI and SME2 to accelerate its speech recognition and language-model components.
Learn how to build a customer support chatbot for Android using Llama 3.2, ExecuTorch, and KleidiAI to run on-device inference on Arm platforms.
Evaluation Conditions
Text Generation · vivo X300 · INT4 · MMLU (test)
| Metric | Resultsvs baseline | Change | |
|---|---|---|---|
| End to end latency p50 | 1459.50ms40378.50ms | −38919.00 ms | |
| End to end latency p90 | 1521.00ms43270.00ms | −41749.00 ms | |
| End to end latency p99 | 1540.00ms44084.00ms | −42544.00 ms | |
| Peak memory | 858.38MB2993.56MB | −2135.19 MB | |
| Time to first inference | 1287.00ms44084.00ms | −42797.00 ms | |
| Tokens per second | 100.36tok/s3.25tok/s | +97.11 tok/s | |
| Time to first token | 195.50ms1272.50ms | −1077.00 ms | |
| Prefill / encode time | 195.50ms1272.50ms | −1077.00 ms | |
| MMLU (0-shot) | 31.72%32.56% | −0.84% |
The Optimized Model is provided as a reference implementation to demonstrate and evaluate execution and performance on Arm-based systems. It is not a production-ready or supported solution.
Arm’s publication of the Optimized Model does not constitute an endorsement or certification of the Original Model or a representation that the Optimized Model is suitable for production use or any particular purpose.
To the fullest extent permitted by applicable law (i) the Optimized Model is provided “as is.” Arm makes no representations or warranties that the Original Model, the Optimized Model or their outputs are accurate, safe, secure, non-infringing, legally compliant, suitable for production use or fit for any particular purpose; and (ii) Arm will not be liable for any loss or damage arising from or in connection with the Optimized Model, its use or its outputs.
You are responsible for independently evaluating the Optimized Model, its outputs and its suitability for your intended use, including compliance with applicable legal, regulatory, safety and security requirements.
Arm does not commit to provide ongoing support, maintenance or updates for the Optimized Model. Any use of or reliance on the Optimized Model or its outputs is at your own risk.
Use of the Portal is subject to the Arm Website Terms and Conditions of Use.
| Configuration | Baseline | Arm Optimized |
|---|---|---|
| Dataset | MMLU (test) | MMLU (test) |
| Sample count | 14,042 | 14,042 |
| Batch size | 1 | 1 |
| Prompt length | 125 | 125 |
| Generation length | 128 | 128 |
| Runs | 20 | 20 |
| Warmup runs | 5 | 5 |