As of 1.0, the monolithic flutter_gemma plugin is split into a small
core package plus opt-in packages for each engine / backend. Your app
ships only the native weight it actually uses. All packages live in one monorepo
(a Dart pub workspace) and the opt-in packages depend on core one-directionally.
The packages#
| Package | What it does | Platforms |
|---|---|---|
flutter_gemma |
Core — registry, contracts, model management, sessions, chat. No engine on its own. Always required. | All |
flutter_gemma_litertlm |
.litertlm
inference via
dart:ffi
(LiteRT-LM C API). Owns the shared native library.
|
Mobile + Desktop + Web |
flutter_gemma_mediapipe |
.task / .bin inference via MediaPipe. |
Mobile + Web |
flutter_gemma_builtin_ai |
System OS models — Gemini Nano (Android / AICore), Apple Foundation Models (iOS 26+/macOS), and Gemini Nano via the Chrome Prompt API (Web). No model file to bundle or download. | Android + iOS + macOS + Web |
flutter_gemma_onnx |
Text generation (
OnnxEngine
) + embeddings (
OnnxEmbeddingBackend
) — ORT-GenAI/ORT via
dart:ffi
on native, Transformers.js/onnxruntime-web on Web.
|
macOS, Linux, Windows, Android, iOS (arm64) + Web |
flutter_gemma_embeddings |
Runtime-agnostic text-embedding pipeline (tokenizer, pooling, isolate worker). Needs a backend —
LiteRtEmbeddingBackend
(
flutter_gemma_litertlm
) or
OnnxEmbeddingBackend
(
flutter_gemma_onnx
).
|
All |
flutter_gemma_rag_qdrant |
On-device RAG vector store (qdrant-edge, via the official qdrant_edge UniFFI SDK). Fastest on native. | Native (no Web) |
flutter_gemma_rag_sqlite |
On-device RAG vector store — in-SQLite KNN via the
sqlite-vec
(
vec0
) extension. Exact + portable.
|
All (incl. Web) |
flutter_gemma_agent |
On-device agent skills — SKILL.md catalog + tool-calling loop (text / JS / native-intent / MCP). | Native, no Web (JS skills: no Linux) |
flutter_gemma_speech |
On-device
speech
— speech-to-text + text-to-speech + a
VoiceSession
voice loop (moonshine/Whisper/Parakeet STT + Matcha/Qwen3/Inflect TTS) via the LiteRT C API +
dart:ffi
.
|
Native (no Web) |
How it works#
-
Core registers no engine by itself. You wire the packages you added through
FlutterGemma.initialize(inferenceEngines:, embeddingBackends:, vectorStore:). See Installation. -
Probe-chain registry. Engines and backends are pure factories that declare
canHandle(spec)+ a priority. The registry selects a provider per model by declaredModelFileType—task/binary→ MediaPipe,litertlm→ LiteRT-LM,onnx→ Onnx,builtIn→ BuiltInAi. -
One app can run both formats. Register both
LiteRtLmEngine()andMediaPipeEngine(), and the registry routes each model to the engine that handles its declaredModelFileType— not its file extension. -
Shared native library.
flutter_gemma_litertlmowns the native LiteRT library (fetched at build time via its Native-Assets hook);flutter_gemma_embeddingsandflutter_gemma_speechhave no hook of their own and consume that bundle transitively.flutter_gemma_onnxowns its own separate ORT / ORT-GenAI native archives.
Choosing packages#
| You want to… | Add |
|---|---|
Run .litertlm models (Gemma 4, Qwen3, FastVLM, + all desktop) |
flutter_gemma_litertlm |
Run .task / .bin models (Gemma3n, Gemma 3, DeepSeek, Qwen 2.5, Phi-4) |
flutter_gemma_mediapipe |
| Run the OS system model with no download (Gemini Nano / Apple Foundation Models) | flutter_gemma_builtin_ai |
| Run ONNX models — ORT-GenAI (native) or Transformers.js (Web) | flutter_gemma_onnx |
| Generate text embeddings |
flutter_gemma_embeddings
+
flutter_gemma_litertlm
(
LiteRtEmbeddingBackend
)
|
| Generate text embeddings from ONNX/ORT models |
flutter_gemma_embeddings
+
flutter_gemma_onnx
(
OnnxEmbeddingBackend
)
|
| On-device RAG on native, fastest (Android/iOS/desktop) | flutter_gemma_rag_qdrant |
| On-device RAG on web, or a portable/exact store on any platform | flutter_gemma_rag_sqlite |
| On-device agent skills the model runs itself (text / JS / native-intent / MCP) | flutter_gemma_agent |
| Transcribe audio, synthesize speech, or run a voice loop on-device (STT + TTS + voice) | flutter_gemma_speech |
Migrating from the 0.16.x monolith is just adding these packages plus one
initialize(...) call — every model / session / chat / embedding / RAG API is
unchanged. See Migration (0.x → 1.0).
ONNX Runtime engine#
flutter_gemma_onnx adds a second inference/embedding stack alongside
LiteRT-LM and MediaPipe: ONNX Runtime — dart:ffi on native (no JVM, no
gRPC), Transformers.js / onnxruntime-web on Web. Two independent arms, either
can be registered on its own:
-
OnnxEngine— text generation via ORT-GenAI on native, Transformers.js on Web. Text-only, greedy decoding, one session at a time in v1 — no vision, no audio, no LoRA yet. -
OnnxEmbeddingBackend— embeddings via plain ONNX Runtime (WordPiece/BERT-style and SentencePiece models) on native, onnxruntime-web on Web, priority 10 overLiteRtEmbeddingBackend's catch-all priority 0.
import 'package:flutter_gemma/flutter_gemma.dart';
import 'package:flutter_gemma_onnx/flutter_gemma_onnx.dart';
await FlutterGemma.initialize(
inferenceEngines: [OnnxEngine()],
embeddingBackends: [OnnxEmbeddingBackend()],
);
Platform support: device-verified via dart:ffi on macOS (Apple Silicon),
Linux x64, Windows x64, Android (arm64), and iOS (arm64). On Web,
OnnxEngine runs generation through Transformers.js v4 and
OnnxEmbeddingBackend runs through onnxruntime-web (WebGPU/WASM) — same
public API as native, no if (kIsWeb) needed in app code.
See the flutter_gemma_onnx README
for the full platform matrix and FFI details.
Genkit integration#
Two companion packages integrate flutter_gemma with Genkit, Google's framework for building AI features:
| Package | What it does | Depends on |
|---|---|---|
genkit_flutter_gemma |
Exposes flutter_gemma as a Genkit model/embedder provider — call
ai.generate(model: flutterGemma.model(...))
and
ai.embed(...)
through the standard Genkit API.
|
flutter_gemma + genkit |
genkit_hybrid |
Provider-agnostic hybrid routing: combine an on-device and a cloud model behind one routing policy, with correct streaming + before-first-token fallback. | genkit only (no flutter_gemma) |
See Genkit for setup and examples.