AI PORTAL
  • Models
  • Code
  • Tools

About the models on this Arm AI Portal

The Arm AI Portal (the “Portal”) provides details of AI models developed, optimized or tested for use on Arm-based platforms. Models listed on the Portal may have been developed by Arm or by third parties. Unless expressly stated otherwise, Arm did not develop or train the underlying third-party model and is not responsible for its original intended behavior or use.

Information accompanies each model listing and describes the model’s provenance, together with relevant information about the original third-party model, its developer and repository, where applicable. Inclusion on the Portal of a model developed by a third party or based on a third-party model does not constitute Arm’s endorsement or certification of that third-party model or any third-party work relating to it.

Except as expressly stated in the information accompanying a specific model, Arm makes no representation as to the accuracy, safety, security, non-infringement, legal compliance, suitability for production use or fitness for any particular purpose of any model, converted or optimized version, or related outputs.

Inclusion of a model on the Portal does not itself grant you any rights to use that model. Use of each model is subject to the applicable license terms, usage restrictions and documentation. You are responsible for reviewing that information and independently evaluating the model, its outputs and its suitability for your intended use, including compliance with applicable legal, regulatory, safety and security requirements. Any use of or reliance on a model, its outputs or related materials is at your own risk. Where the Portal links to repositories or other materials controlled by third parties, those repositories and materials are outside Arm’s control and may change without notice.

Use of the Portal is subject to the Arm Website Terms and Conditions of Use.

Models

Filters

1

Device Class

Edge LinuxMobile CPUCloud CPU

Evaluation Target

vivo X300AWS Graviton G4Raspberry Pi 5

Arm Technology

SME2

Task

Image ClassificationText GenerationObject DetectionImage SegmentationAutomatic Speech RecognitionFeature ExtractionImage To ImageText To SpeechDepth EstimationImage Text To TextKeypoint DetectionText To ImageTranslationZero Shot Image Classification

Vendor

Runtime

Quantization

Format

Execution Backend

Model Size

≤5.7 GB

Average Memory

≤5.48 GB

Peak Memory

≤5.5 GB
    22 results

    Qwen3-0.6B INT4

    AWS Graviton G4

    A compact decoder-only causal language model for text generation and instruction following, using GPTQ INT4 weight quant

    text-generation483.7 MBONNX

    Gemma-3-1B-Base INT4

    vivo X300SME2

    A compact decoder-only causal language model for text generation, using INT4 quantized linear weights with dynamic INT8

    text-generation865.5 MBONNX

    ERNIE-4.5-0.3B-PT Q4_K_M

    Raspberry Pi 5

    ERNIE-4.5-0.3B-PT optimized for text generation in GGUF format with the llama.cpp runtime, targeting Arm-based Edge Linu

    text-generation240.6 MBllama.cpp

    Llama-3.1-8B-Instruct INT4

    AWS Graviton G4

    Instruction-tuned causal language model for chat and text generation, packaged in ONNX format with INT4 weight quantizat

    text-generation5.4 GBONNX

    Qwen3-1.7B-Base INT4 Dynamic

    vivo X300SME2

    This is a 4-bit quantized version of Qwen/Qwen3-1.7B-Base optimized for edge deployment on ARM devices using ONNX Runtim

    text-generation1.3 GBONNX

    Llama-3.2-1B-Base INT4

    AWS Graviton G4

    Compact decoder-only causal language models optimized for efficient text generation on ARM CPUs, using ONNX Runtime GenA

    text-generation1.2 GBONNX

    Gemma-4-E2B 8da4w

    vivo X300SME2

    Instruction-tuned autoregressive text-generation model compressed with INT4 weights and INT8 activations.

    text-generation2.7 GBExecuTorch

    Qwen3.5-0.8B Q4_K_M

    vivo X300SME2

    This is a 4-bit weight-quantized version of Qwen/Qwen3.5-0.8B packaged as a single GGUF file for edge deployment on ARM

    text-generation529.3 MBllama.cpp

    DeepSeek-R1-Distill-Qwen-1.5B Q4_K_M

    vivo X300SME2

    DeepSeek-R1-Distill-Qwen-1.5B text generation optimized as a Q4_K_M GGUF model for the llama.cpp runtime on Arm-based Pr

    text-generation1.1 GBllama.cpp

    Qwen3-1.7B-Base Q4_K_M

    vivo X300SME2

    This is a 4-bit weight-quantized version of Qwen/Qwen3-1.7B-Base packaged as a single GGUF file for edge deployment on A

    text-generation1.2 GBllama.cpp

    Qwen3-0.6B-Base Q4_K_M

    Raspberry Pi 5

    This is a Q4_K_M k-quantized GGUF build of Qwen/Qwen3-0.6B-Base — the base pretrained checkpoint, not the instruct varia

    text-generation434.4 MBllama.cpp

    Gemma-3-1B-Instruct INT4

    vivo X300SME2

    Instruction-following causal language model packaged for efficient text generation, using INT4 quantized weights, dynami

    text-generation865.5 MBONNX

    TinyLlama-1.1B-Chat INT4

    AWS Graviton G4

    A compact chat-oriented causal language model family optimized with INT4/INT8 quantization for efficient text generation

    text-generation763.2 MBONNX

    Llama-3.1-8B-Base INT4

    AWS Graviton G4

    A quantized 8B-parameter base causal language model for text completion, using INT4 weights with dynamic INT8 activation

    text-generation5.7 GBONNX

    SmolLM2-360M-Instruct 8da4w

    vivo X300SME2

    A compact 362M-parameter decoder-only causal language model for instruction-following, open-ended text generation, and m

    text-generation251.7 MBExecuTorch

    Llama-3.2-1B-Instruct INT4

    AWS Graviton G4

    Instruction-following causal language model family with a decoder-only Transformer architecture, packaged in ONNX format

    text-generation1.2 GBONNX

    Llama-3.2-3B-Instruct INT4

    AWS Graviton G4

    Quantized instruction-following causal language models for efficient text generation, using INT4 weights with dynamic IN

    text-generation2.7 GBONNX

    Llama-3.2-1B-Instruct INT4

    vivo X300SME2

    Quantized decoder-only causal language model for instruction-following and text generation, packaged in ONNX GenAI forma

    text-generation1.2 GBONNX

    TinyLlama-1.1B-Chat INT4

    vivo X300SME2

    A compact chat-oriented causal language model for text generation and instruction following, using mixed-precision quant

    text-generation763.2 MBONNX

    Qwen3.5 2B Q4_K_M

    vivo X300SME2

    Qwen3.5-2B text generation, quantized to Q4_K_M GGUF for the llama.cpp runtime on Arm-based premium smartphone platforms

    text-generation1.3 GBllama.cpp

    Develop on Arm with your AI coding assistant

    The Arm AI Portal MCP Server makes it faster to find optimized models, integrate code examples and deploy to your target device from your agentic coding assistant.

    User guide

    Run the following command in your terminal:

    codex mcp add arm-ai --url https://mcp.api.devplatform.arm.com/ai-portal