Core AI model models tested and mapped to practical AI features to build from:
Realtime video image detection, tracking, and segmentation: SAM 3 & RF-DETR: RealtimeObjectDetection.swift
AI feature
Core AI model candidates
Audio-text-to-text
Qwen2.5-Omni-3B Audio ; pipeline option: Whisper or Wav2Vec 2.0 plus a text-generation LLM such as Qwen3 , Gemma 4 E2B , or Qwen3.5
Image-text-to-text
Qwen3-VL , MiniCPM-V 4.6 , Gemma 4 E2B vision
Image-text-to-image
No matching model found in the referenced Core AI repos
Image-text-to-video
No matching model found in the referenced Core AI repos
Visual question answering
Qwen3-VL , MiniCPM-V 4.6 , Gemma 4 E2B vision
Document question answering
Unlimited-OCR plus a text LLM; Qwen3-VL , MiniCPM-V 4.6 , Gemma 4 E2B vision
Video-text-to-text
No matching model found in the referenced Core AI repos
Visual document retrieval
CLIP , Unlimited-OCR , EmbeddingGemma 300M , Qwen3-Embedding 0.6B , Qwen3-Reranker 0.6B
Any to any
No matching model found in the referenced Core AI repos
AI feature
Core AI model candidates
Depth estimation
Depth Anything v3 , Depth Anything 3
Image classification
PVT v2 , CLIP
Object detection
YOLOS , RF-DETR nano/small/medium/large
Image segmentation
EfficientSAM , SAM 3 , RF-DETR-Seg
Text to Image
Stable Diffusion 1.5, 2.1, 3.5 Medium , FLUX.2
Image to text
Qwen3-VL , MiniCPM-V 4.6 , Gemma 4 E2B vision , Unlimited-OCR
Image to image
EDSR , AdcSR x4
Image to video
No matching model found in the referenced Core AI repos
Unconditional image generation
No matching model found in the referenced Core AI repos
Video classification
No matching model found in the referenced Core AI repos
Text to video
No matching model found in the referenced Core AI repos
Zero-shot image classification
CLIP
Mask generation
EfficientSAM , SAM 3 , RF-DETR-Seg
Zero-shot object detection
No matching model found in the referenced Core AI repos
Text to 3D
No matching model found in the referenced Core AI repos
Image to 3D
No matching model found in the referenced Core AI repos
Image feature extraction
CLIP , PVT v2
Keypoint detection
No matching model found in the referenced Core AI repos
Video to video
No matching model found in the referenced Core AI repos
Natural Language Processing
AI feature
Core AI model candidates
Text classification
RoBERTa , Gemma 3 , GPT-OSS , Mistral , Mixtral , Qwen2.5 , Qwen3 , Qwen3 MoE , Gemma 4 E2B/E4B/12B/31B , Qwen3.5 , Qwen3.6 , GLM-4.7-Flash , LFM2.5 , Granite 4.0-H , Nanbeige4.1-3B
Token classification
RoBERTa
Table question answering
General LLM option: Gemma 3 , GPT-OSS , Mistral , Mixtral , Qwen3 , Gemma 4 , Qwen3.5 , GLM-4.7-Flash , Granite 4.0-H
Question answering
T5 , Gemma 3 , GPT-OSS , Mistral , Mixtral , Qwen2.5 , Qwen3 , Qwen3 MoE , Gemma 4 , Qwen3.5 , Qwen3.6 , GLM-4.7-Flash , LFM2.5 , Granite 4.0-H , Nanbeige4.1-3B
Zero-shot classification
Gemma 3 , GPT-OSS , Mistral , Mixtral , Qwen3 , Gemma 4 , Qwen3.5 , GLM-4.7-Flash , Granite 4.0-H
Translation
T5 , Gemma 3 , GPT-OSS , Mistral , Mixtral , Qwen2.5 , Qwen3 , Qwen3.5 , Qwen3.6 , Gemma 4 , LFM2.5
Summarization
T5 , Gemma 3 , GPT-OSS , Mistral , Mixtral , Qwen2.5 , Qwen3 , Qwen3 MoE , Gemma 4 , Qwen3.5 , Qwen3.6 , GLM-4.7-Flash , LFM2.5 , Granite 4.0-H
Feature extraction
RoBERTa , T5 , EmbeddingGemma 300M , Qwen3-Embedding 0.6B
Text generation
Gemma 3 , GPT-OSS , Mistral , Mixtral , Qwen2.5 , Qwen3 , Qwen3 MoE , Gemma 4 E2B , Gemma 4 E4B , Gemma 4 12B , Gemma 4 31B , Qwen3.5 , Qwen3.6-35B-A3B , Qwen3.6-27B , GLM-4.7-Flash , LFM2.5-1.2B , LFM2.5-8B-A1B , Granite 4.0-H , Nanbeige4.1-3B
Fill-mask
RoBERTa
Sentence similarity
EmbeddingGemma 300M , Qwen3-Embedding 0.6B , RoBERTa
Text ranking
Qwen3-Reranker 0.6B , EmbeddingGemma 300M , Qwen3-Embedding 0.6B
AI feature
Core AI model candidates
Tabular classification
No matching model found in the referenced Core AI repos
Tabular regression
No matching model found in the referenced Core AI repos
Time series forecasting
No matching model found in the referenced Core AI repos