HomeCommunityAI blog
October 1, 2026

On-device toxic speech detection with ML and SME2

See how LiteRT, quantization, KleidiAI, and SME2 improve on-device toxic speech detection on mobile, reducing latency and server processing demand overall

By Ben Clark

Share
Reading time 8 minutes

Many games and apps have voice streaming capability alongside the core play experience. This enables people to enjoy socializing during a shared activity. But how can we make sure that people do not abuse that facility to harass and abuse others?

Roblox take this seriously, as they must with their core young audience. The company developed a machine learning (ML) network to detect toxic speech and provide voice safety. The model detects profanity, harassment, discrimination, sexual language, and other categories of speech that Roblox considers unacceptable for their platform.

Currently this network would need to run on servers and process a vast number of voice channels. This has a significant cost, whether it runs on a company’s own servers or in the cloud. At Arm we asked whether the network could run on device, fast enough, and without impacting game performance?

As a side note, on-device moderation cannot be treated as sole proof that audio is safe. A modified client can alter the results. In production, local inference should be an optimization rather than the only safeguard. Some server-side auditing and enforcement will still be needed, with a fallback for unsupported or untrusted clients, if those clients are not banned.

Initial investigations

I started by converting the network to an ML runtime that is efficient on mobile. This involved some challenges and required a compromise. The compromise is necessary because of the limited support for dynamic input for mobile ML runtimes.

Dynamic input is needed because the length of speech varies. Processing the full 15s maximum chunk length of Roblox’s model would otherwise waste resources. ExecuTorch supports some dynamic input, so I started with that runtime. However, LiteRT delivered better performance for this network for the full 15s than ExecuTorch did for anything more than 1s long. Different networks perform better on different backends, so it was wise to try both. Arm is also working on an ExecuTorch improvement that is expected to improve performance for this network by 3x.

Performance with full float32 was not good enough, so I investigated quantization. This highlighted another issue: ExecuTorch dynamic support is not compatible with quantized networks. The quantized networks ran much faster on ExecuTorch than the float32, but still not as fast as LiteRT. I therefore decided to continue with LiteRT and make versions of the network for a few different lengths, such as 5s, 10s, and 15s. The application can then use the appropriate network as needed.

The CPU was the expected main target because games typically use the GPU for graphics. However, I also wanted to check GPU performance. This created both an opportunity and a challenge. To run the network on GPU, I needed to modify the network significantly. In doing so I also improved performance.

First, I could pre-computed the start of the network because the input was no longer dynamic. Next, the GPU does not support GATHER_ND operator. However, the operators were not needed, so I replaced the GroupNorm implementation with a more efficient version.

The final significant adjustment was to the WavLM attention. At the time, the LiteRT GPU backend did not support 5D tensors. I therefore modified the network to remove the 5D tensors and allow it to run on GPU. Support for 5D tensors was since added in v2.1.6

The network ran quite quickly on GPU with fp16, before considering further quantization. Disappointingly though, it also found a bug in the LiteRT GPU backend that badly affected the output values. The LiteRT team continues to investigate the issue. For now, the investigation is only focused on the CPU.

CPU optimization and performance

So, what performance can we achieve on CPU, and are there quality implications?

Our initial unquantized model, uses only one thread, leaving the rest of the mobile resources for the game. With LiteRT, the 15s version of the network takes 2s to run on a phone with a MediaTek Dimensity 9500 chip containing Arm C1 cores. Quality is identical to the model before conversion from PyTorch to tflite format, but latency certainly could be improved.

How can we speed this up? ML models are mainly matrix multiplications, and there are Arm technologies that make these faster. Neon and SVE enable vector calculations to be parallelized. Smaller data types, such as 8-bit integer (int8), fit more into the vector and do more at once.

Scalable Matrix Extension 2 (SME2), enables calculations across a whole matrix of numbers. It has been available in Mac, Android, and iOS devices since the end of 2025, with more phones adding support this year.

To support SME2 adoption, Arm has released KleidiAI. Kleidi AI is integrated into ML inference frameworks, so developers get automatic access to Neon/SVE2/SME2 kernels. By optimizing our model for these technologies, we can achieve significant performance improvements.

The first technique to try is dynamic quantization. The weights are quantized to int8, but the activations are left unquantized. The model size is mostly weights, so it shrinks to nearly a quarter of its original size. Google’s ai_edge_quantizer makes this relatively straightforward after converting from PyTorch to tflite.

Performance improves to just over 0.6s on 1 thread on our test phone for 15s of input. In the initial test, there was no SME2 speed-up for int8 dynamic input because KleidiAI kernels had not been added for it yet. In recent re-testing, newly added kernals improved performance by about 10%, reducing latency to less than 0.54s.

More SME2 dynamic int8 kernels continue to be added to KleidiAI. These kernels then become available through XNNPACK, which LiteRT and ExecuTorch use as their main CPU backend.

This is good, but can we do better? KleidiAI and XNNPACK provide better performance support for static quantization. This could reduce latency if the model retains its quality.

I created a test set of more than 40 spoken phrases. Some were clearly toxic, some were benign controls, and some that tried to find the middle ground. I kept my headphones on so that I did not get complaints from my colleagues about all the swearing!

The model outputs six values that each correspond to a probability that the speech processed contains toxic language in a particular category. The categories include profanity, harassment, and sexual. Dynamic quantization produced slightly different output values for many phrases. However, the values remained similar enough that the results would hold. Static quantization proved more challenging.

The default settings lost a lot of signal, and further investigation was needed. ai_edge_quantizer offers several options, but the default uses a naïve minimum and maximum algorithm over your representative dataset. Other tools often provide only this option. With ai_edge_quantizer, I could control the algorithm and choose which layers and operators were quantized. I did not want to quantize and dequantize repeatedly as that would lose my performance gains. I therefore needed to find a delicate balance.

In the end I got good results with:

  • Excluding the first and last 3 layers, which are more sensitive to quantization.
  • Using the mean squared error (MSE) quantization algorithm rather than the naïve min/max.

MSE quantizes only convolutions and fully connected layers. This leaves many quantization and dequantization pairs around those operators, but still a net performance gain. I improved further by adding dynamic quantization to batch matrix multiplies (BMMs). BMM is the only additional operator dynamic quantization works on. Quantization tools still need significant further development to take what is theoretically possible in academic research into actual production.

With these values I get great performance, but before we reveal that, does the quality hold up? The answer is mostly yes. Outputs with a high toxicity probability, above 0.5 in any category, remained high. Low probability outputs also remained low. However, outputs that originally fell between 0.25 and 0.5 decreased significantly after quantization. An example of this is "wow you are a stupid idiot," which originally produced a harassment score of 0.42. I believe this could be fixed with quantization-aware training (QAT) of the network. However, without the full training dataset this cannot be tested.

As to the performance, with SME2 used on our test phone it takes 0.37s to process 15s of audio on 1 thread. Or 0.096s for the 5s version of the network. SME2 continues to be an important part of the performance here. Without SME2, the 15s version takes 0.59s. With SME2, it runs about 60% faster. The 5s version achieves a 70% speed-up.

Conclusion

These results show that on-device toxic speech filtering is feasible on current SME2-enabled mobile devices. Using just one CPU thread, the model processes 15s of audio in 0.37s, or 0.096s for 5s. This leaves useful CPU headroom for a game while reducing the need to send every audio segment to a server.

Quantization made this performance practical, while SME2 provided a substantial further improvement on supported devices. Not all devices will be able to run the model efficiently. A production deployment would need to decide whether target hardware is suitable, and if the client is trustworthy.

Even if some devices cannot run the filter, on-device inference could reduce server processing costs for game and app companies. It could also provide faster responses by avoiding network latency.

The next step would be using quantization-aware training on the model with full training data. We would also need to test all supported languages and the six toxicity categories across more devices.

If you are interested in exploring further, you can learn more about KleidiAi in this blog post. You can also follow our Learning Paths to understand how to Accelerate LiteRT Models on Android with KleidiAI and SME2 and Profile the Performance of AI and ML Mobile Applications on Arm.

Accelerate LiteRT Models on Android with KleidiAI and SME2   Profile the Performance of AI and ML Mobile Applications on Arm


2Log in to like this post
Share

Article text

Re-use is only permitted for informational and non-commercial or personal use only.

placeholder