Genkit is Google's open-source framework
for building AI-powered features in Dart and Flutter. Two packages bridge
flutter_gemma into Genkit — one wraps the on-device runtime as a standard
Genkit provider, the other adds hybrid routing so you can combine on-device
and cloud models behind a single ai.generate call.
genkit_flutter_gemma#
Wraps flutter_gemma as a Genkit model and embedder provider. Once registered, every Genkit feature (streaming, tool use, embeddings, prompt templates) works with the on-device model exactly as it would with any cloud provider.
Add to pubspec.yaml#
dependencies:
genkit: ^0.16.0 # the framework itself — every snippet below uses it
genkit_flutter_gemma: ^0.6.1
flutter_gemma: ^1.8.3
# Add the inference engine(s) you need:
flutter_gemma_litertlm: ^1.7.0 # .litertlm models (mobile + desktop + web) + LiteRtEmbeddingBackend
flutter_gemma_mediapipe: ^1.0.6 # .task / .bin models (mobile + web)
# Optional — for embeddings (needs a backend, e.g. flutter_gemma_litertlm above):
flutter_gemma_embeddings: ^2.1.1
Setup#
Register the engine packages in FlutterGemma.initialize(), install your
model, then create a Genkit instance with the plugin:
import 'package:flutter_gemma/flutter_gemma.dart';
import 'package:flutter_gemma_litertlm/flutter_gemma_litertlm.dart';
import 'package:flutter_gemma_mediapipe/flutter_gemma_mediapipe.dart';
import 'package:genkit/genkit.dart';
import 'package:genkit_flutter_gemma/genkit_flutter_gemma.dart';
// 1. Register providers (call once in main).
await FlutterGemma.initialize(
inferenceEngines: const [LiteRtLmEngine(), MediaPipeEngine()],
embeddingBackends: const [LiteRtEmbeddingBackend()], // flutter_gemma_litertlm
);
// 2. Install the model (host app responsibility).
await FlutterGemma.installModel(modelType: ModelType.gemmaIt)
.fromAsset('assets/gemma-3-1b-it-int4.task')
.install();
// For a .litertlm model declare the type in BOTH places — installModel(
// fileType: ModelFileType.litertlm) and FlutterGemmaModelConfig(fileType: ...).
// Both default to ModelFileType.task, which routes the model to MediaPipe.
// 3. Create a Genkit instance with the plugin.
final ai = Genkit(plugins: [
GenkitFlutterGemmaPlugin(
models: [
FlutterGemmaModelConfig(
name: 'gemma-3-nano',
modelType: ModelType.gemmaIt,
),
],
embedders: [
FlutterGemmaEmbedderConfig(name: 'embedding-gemma-300m'),
],
),
]);
Generate text#
final response = await ai.generate(
model: flutterGemma.model('gemma-3-nano'),
prompt: 'Hello!',
);
print(response.text);
Stream text#
final stream = ai.generateStream(
model: flutterGemma.model('gemma-3-nano'),
prompt: 'Write a short story.',
);
await for (final chunk in stream) {
stdout.write(chunk.text);
}
Embeddings#
final embeddings = await ai.embed(
embedder: flutterGemma.embedder('embedding-gemma-300m'),
documents: [
DocumentData(content: [TextPart(text: 'Flutter is a UI toolkit.')]),
],
);
Configuration options#
Pass FlutterGemmaModelOptions to tune inference:
final response = await ai.generate(
model: flutterGemma.model('gemma-3-nano'),
prompt: 'Hello!',
config: FlutterGemmaModelOptions(
maxTokens: 2048,
temperature: 0.5,
topK: 40,
supportImage: true,
toolChoice: 'auto', // tool calling mode: 'auto' / 'required' / 'none'
// Optional per-component backend ('cpu'/'gpu'/'npu'):
preferredBackend: 'gpu', // text decoder
preferredAudioBackend: 'gpu', // audio encoder (~2x on Metal; defaults to CPU)
// preferredVisionBackend defaults to CPU (Metal/WebGPU can't run its ops).
),
);
Prefer Genkit's standard top-level parameter — ai.generate(toolChoice: 'none')
— which takes precedence over the toolChoice config field above (kept as a
legacy fallback). Either way: 'auto' lets the model decide, 'required'
forces
a tool call, 'none' forbids one. An unrecognized value throws
INVALID_ARGUMENT rather than quietly falling back to 'auto'.
Structured (JSON) output#
The plugin advertises output: ['text', 'json']. On-device Gemma has no native
schema-constrained decoder, so Genkit's instruction-injection fallback drives
JSON output — the plugin returns raw model text and Genkit's extractJson
populates response.output. Pass an outputSchema and read the parsed object:
final response = await ai.generate(
model: flutterGemma.model('gemma-3-nano'),
prompt: 'Give me a pancake recipe.',
outputSchema: Recipe.$schema, // any @Schema()-annotated type
);
final Recipe? recipe = response.output;
Context-window trimming#
On-device models run with a fixed, small context window (maxTokens — 1024 for
most .litertlm models). A long multi-turn chat overflows it and the native
runtime fails to allocate the KV cache mid-generation. trimContext() is a
middleware that drops the oldest non-system turns before each model call,
always keeping every system message and the most recent message:
final response = await ai.generate(
model: flutterGemma.model('gemma-3-nano'),
prompt: 'Continue our conversation…',
messages: longHistory,
use: [trimContext(maxInputTokens: 800)],
);
With no arguments the budget is derived from the request's maxTokens (the
model's context window) minus 256 tokens of response headroom.
genkit_hybrid#
Provider-agnostic hybrid routing for Genkit. Combine any two existing Genkit
models — on-device, cloud, or anything else — behind one routing policy. The
result is an ordinary Model, so your app still calls a single ai.generate.
genkit_hybrid depends only on genkit — it has no dependency on
flutter_gemma and works with any pair of Genkit models.
Add to pubspec.yaml#
dependencies:
genkit_hybrid: ^0.2.1
genkit: ^0.16.0
Basic usage#
import 'package:genkit/genkit.dart';
import 'package:genkit_hybrid/genkit_hybrid.dart';
final ai = Genkit();
// onDeviceModel and cloudModel are ordinary Genkit Models you already have.
final smart = hybridModelOnDeviceCloud(
onDevice: onDeviceModel,
cloud: cloudModel,
strategy: ConnectivityStrategy(
isOnline: () => connectivity.isOnline,
online: kCloud,
offline: kOnDevice,
),
);
// A hybrid model is an ordinary Model — register it, then use it like any other.
ai.registry.register(smart);
final response = await ai.generate(model: smart, prompt: 'Hello!');
Routing strategies#
| Strategy | Routes on |
|---|---|
PreRoutingStrategy(fn) | your own function (privacy, cost, user tier…) |
FallbackStrategy(order) |
fixed priority order — kOnDevice first or kCloud first |
ConnectivityStrategy(...) | network availability |
InputSizeStrategy(...) | prompt length |
CapabilityStrategy(supports: {...}) |
the capabilities a request needs — vision / audio / tools / json — routing to branches that declare them |
CostStrategy(budgetAvailable:, premium:, cheap:) |
a budget signal — the premium branch while the budget holds, the cheap branch once it's spent |
FirstMatch([...]) | first child strategy that decides (chain of rules) |
WithFallback(s, fallbackOrder: order) |
any strategy's pick + a guaranteed fallback tail |
Prefer on-device, fall back to cloud#
hybridModelOnDeviceCloud(
onDevice: onDeviceModel,
cloud: cloudModel,
strategy: FallbackStrategy([kOnDevice, kCloud]),
);
Chain multiple rules#
hybridModelOnDeviceCloud(
onDevice: onDeviceModel,
cloud: cloudModel,
strategy: WithFallback(
FirstMatch([
PreRoutingStrategy((c) => userOptedOutOfCloud ? kOnDevice : ''),
ConnectivityStrategy(
isOnline: () => net.isOnline,
online: kCloud,
offline: kOnDevice,
),
]),
fallbackOrder: [kOnDevice],
),
);
Route by required capabilities#
CapabilityStrategy inspects what the request actually needs — an image or
audio part, tool definitions, or JSON output — and keeps only the branches that
declare those capabilities (or [] when none qualifies, so compose it with
WithFallback). Capabilities are declared explicitly per branch; nothing is
inferred from model metadata.
final smart = hybridModel(
branches: {'onDevice': onDeviceModel, 'cloud': cloudModel},
strategy: WithFallback(
CapabilityStrategy(supports: {
'onDevice': {}, // text only
'cloud': {ModelCapability.vision, ModelCapability.tools},
}),
fallbackOrder: ['cloud'],
),
);
A plain-text request can use either branch; a request carrying an image or tool
definitions is routed to cloud, the only branch that declares those
capabilities.
Budget-gate a paid branch#
CostStrategy sends traffic to a premium branch only while an app-supplied
budget signal holds, and falls back to the cheap branch once it's spent. Your
app owns the accounting (running spend, a daily cap, a quota) and reduces it to
one bool — the package depends on no billing SDK.
hybridModel(
branches: {'onDevice': onDeviceModel, 'cloud': cloudModel},
strategy: CostStrategy(
budgetAvailable: () => spend.today < dailyCap,
premium: 'cloud',
cheap: 'onDevice',
),
);
Error policy and fallback#
Fallback is error-driven: the strategy picks an order, and the next branch is
tried only when the current one throws, and only on a transient failure —
any non-GenkitException error (network, timeout, OOM), or a GenkitException
with UNAVAILABLE, DEADLINE_EXCEEDED, RESOURCE_EXHAUSTED
or INTERNAL.
Permanent errors — INVALID_ARGUMENT, PERMISSION_DENIED,
UNAUTHENTICATED,
FAILED_PRECONDITION, NOT_FOUND — propagate immediately, since they would
fail the same way on every branch. A GenkitException thrown without an
explicit status defaults to INTERNAL, so it is retried.
During streaming the same policy applies plus a hard cut-off: fallback is possible only before the first token. Once a branch has emitted a chunk, any later failure propagates — a partially delivered response cannot be silently re-routed.
Escalate on a quality check with cascadeModel#
cascadeModel is a Model (not a strategy): it runs branches in order and
escalates to the next one only when your accept predicate rejects the
response — "try the cheap on-device model; go to the cloud only if the answer
isn't good enough". accept is any check you like (a length or regex test, a
JSON-parses check) and may be async, e.g. an LLM-as-judge.
import 'package:genkit_hybrid/genkit_hybrid.dart';
final smart = cascadeModel(
branches: {'onDevice': onDeviceModel, 'cloud': cloudModel},
order: ['onDevice', 'cloud'],
accept: (r) => r.text.trim().length > 20,
);
final response = await ai.generate(model: smart, prompt: 'Explain quantum tunnelling.');
cascadeModelis non-streaming in v1. A quality verdict needs the whole response, and a streamed response can't be un-sent, so a streaming request is run non-streamed and the accepted response is emitted as a single final chunk.