AI PORTAL
  • Models
  • Code
  • Tools

About the models on this Arm AI Portal

The Arm AI Portal (the “Portal”) provides details of AI models developed, optimized or tested for use on Arm-based platforms. Models listed on the Portal may have been developed by Arm or by third parties. Unless expressly stated otherwise, Arm did not develop or train the underlying third-party model and is not responsible for its original intended behavior or use.

Information accompanies each model listing and describes the model’s provenance, together with relevant information about the original third-party model, its developer and repository, where applicable. Inclusion on the Portal of a model developed by a third party or based on a third-party model does not constitute Arm’s endorsement or certification of that third-party model or any third-party work relating to it.

Except as expressly stated in the information accompanying a specific model, Arm makes no representation as to the accuracy, safety, security, non-infringement, legal compliance, suitability for production use or fitness for any particular purpose of any model, converted or optimized version, or related outputs.

Inclusion of a model on the Portal does not itself grant you any rights to use that model. Use of each model is subject to the applicable license terms, usage restrictions and documentation. You are responsible for reviewing that information and independently evaluating the model, its outputs and its suitability for your intended use, including compliance with applicable legal, regulatory, safety and security requirements. Any use of or reliance on a model, its outputs or related materials is at your own risk. Where the Portal links to repositories or other materials controlled by third parties, those repositories and materials are outside Arm’s control and may change without notice.

Use of the Portal is subject to the Arm Website Terms and Conditions of Use.

Models

Develop on Arm with your AI coding assistant

Arm MCP Server for AI Portal makes it faster to find optimized models, integrate code examples and deploy to your target device from your agentic coding assistant.

User guide

Run the following command in your terminal:

codex mcp add arm-ai --url https://mcp.api.devplatform.arm.com/ai-portal

Filters


Device Class

Edge LinuxEthos-U NPUMobile CPUCloud CPU

Evaluation Target

vivo X300Raspberry Pi 5AWS Graviton G4Alif DK-E8

Arm Technology

SME2NX

Task

Image ClassificationText GenerationObject DetectionImage SegmentationAutomatic Speech RecognitionFeature ExtractionImage To ImageText To SpeechDepth EstimationImage Text To TextKeypoint DetectionText To ImageTranslationZero Shot Image Classification

Vendor

Runtime

Quantization

Model Size

≤6 GB

Average Memory

≤8.13 GB

Peak Memory

≤15.37 GB
    84 results

    Qwen3-0.6B INT4

    AWS Graviton G4

    A compact decoder-only causal language model for text generation and instruction following, using GPTQ INT4 weight quant

    text-generation483.7 MBonnx

    Gemma-3-1B-Base INT4

    vivo X300SME2

    A compact decoder-only causal language model for text generation, using INT4 quantized linear weights with dynamic INT8

    text-generation865.5 MBonnx

    ResNet-50 INT8 (per-tensor)

    Alif DK-E8

    A 50-layer residual convolutional neural network for 1000-class image classification, using INT8 per-tensor quantization

    image-classification15.1 MBexecutorch

    DDRNet-23-Slim INT8

    Raspberry Pi 5

    DDRNet-23-Slim semantic segmentation optimized using static INT8 post-training quantization (PTQ) with symmetric per-cha

    image-segmentation5.9 MBexecutorch

    TinySD INT8

    vivo X300SME2

    INT8-quantized text-to-image diffusion models based on Stable Diffusion 1.5, generating 512×512 images from prompts with

    text-to-image1 GBexecutorch

    YOLO11n-Pose INT8

    Raspberry Pi 5

    This is an INT8-quantized version of [yolo11n-pose](https://github.com/ultralytics/ultralytics) (Ultralytics) optimized

    keypoint-detection5.4 MBexecutorch

    ERNIE-4.5-0.3B-PT Q4_K_M

    Raspberry Pi 5

    ERNIE-4.5-0.3B-PT optimized for text generation in GGUF format with the llama.cpp runtime, targeting Arm-based Edge Linu

    text-generation240.6 MBgguf

    Llama-3.1-8B-Instruct INT4

    AWS Graviton G4

    Instruction-tuned causal language model for chat and text generation, packaged in ONNX format with INT4 weight quantizat

    text-generation5.4 GBonnx

    Qwen3-1.7B-Base INT4 Dynamic

    vivo X300SME2

    This is a 4-bit quantized version of [Qwen/Qwen3-1.7B-Base](https://huggingface.co/Qwen/Qwen3-1.7B-Base) optimized for e

    text-generation1.3 GBonnx

    MobileSAM INT8

    Raspberry Pi 5

    An INT8-quantized, lightweight segment-anything image segmentation family with a compact TinyViT backbone, designed for

    image-segmentation22.8 MBexecutorch

    bge-base-en-v1.5 INT8

    vivo X300SME2

    A compact INT8-quantized English text embedding model for feature extraction, semantic search, retrieval, and similarity

    feature-extraction111.2 MBtflite

    realesr-general-x4v3 INT8

    vivo X300SME2

    This is an INT8-quantized version of [realesr-general-x4v3](https://github.com/xinntao/Real-ESRGAN), a blind 4x super-re

    image-to-image1.4 MBexecutorch

    ESRGAN x4 INT8

    AWS Graviton G4

    INT8-quantized 4× single-image super-resolution models for enhancing low-resolution RGB images, especially urban scenes

    image-to-image24.3 MBexecutorch

    YOLOv9-S INT8

    Raspberry Pi 5

    An INT8-quantized single-stage object detection family for efficient edge inference on ARM devices. It targets 80-class

    object-detection12.4 MBexecutorch

    GoogLeNet INT8

    Raspberry Pi 5

    An INT8-quantized Inception v1 image-classification model with per-channel weights and per-tensor activations. It accept

    image-classification6.8 MBexecutorch

    Llama-3.2-1B-Base INT4

    AWS Graviton G4

    Compact decoder-only causal language models optimized for efficient text generation on ARM CPUs, using ONNX Runtime GenA

    text-generation1.2 GBonnx

    Whisper Medium INT8

    vivo X300SME2

    INT8-quantized automatic speech recognition family for English speech-to-text, based on an encoder-decoder Transformer a

    automatic-speech-recognition779.2 MBtflite

    Gemma-4-E2B 8da4w

    vivo X300SME2

    Instruction-tuned autoregressive text-generation model compressed with INT4 weights and INT8 activations.

    text-generation2.7 GBexecutorch

    Swin Tiny INT8

    Raspberry Pi 5

    An INT8-quantized hierarchical Vision Transformer for 1000-class image classification, using shifted-window attention fo

    image-classification29.9 MBexecutorch

    YOLO26n FP16

    vivo X300SME2

    Lightweight object detector for COCO classes, packaged with FP16 weights for efficient edge AI workloads.

    object-detection10.9 MBtflite