Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ and 96% with Gemini 3.1 Pro. For more details, see our blog post:
| [`gemini-api-dev`](skills/gemini-api-dev) | Skill for developing Gemini-powered apps. Provides the best practices for building apps that use the Gemini API. |
| [`gemini-live-api-dev`](skills/gemini-live-api-dev) | Skill for building real-time, bidirectional streaming apps with the Gemini Live API. Covers WebSocket-based audio/video/text streaming, voice activity detection, native audio features, function calling, and session management. |
| [`gemini-interactions-api`](skills/gemini-interactions-api) | Skill for building apps with the [Gemini Interactions API](https://ai.google.dev/gemini-api/docs/interactions?ua=chat). Covers text generation, multi-turn chat, streaming, function calling, structured output, image generation, Deep Research agents, deprecated model guardrails, and both Python and TypeScript SDKs. |
| [`gemini-nano-dev`](skills/gemini-nano-dev) | Skill for building web apps and Chrome Extensions with Chrome's built-in AI powered by Gemini Nano. Covers the Prompt API (LanguageModel), Summarizer, Writer, Rewriter, Proofreader, Language Detector, and Translator APIs for client-side, on-device AI inference. |

## Installation

Expand Down
380 changes: 380 additions & 0 deletions skills/gemini-nano-dev/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,380 @@
---
name: gemini-nano-dev
description: Use this skill when building web applications or Chrome Extensions that use Chrome's built-in AI powered by Gemini Nano. Covers the Prompt API (LanguageModel), Summarizer API, Writer API, Rewriter API, Proofreader API, Language Detector API, and Translator API for client-side, on-device AI inference with no server required.
---

# Gemini Nano (Chrome Built-in AI) Development Skill

## Overview

Chrome's built-in AI APIs let you perform AI-powered tasks directly in the browser using the **Gemini Nano** model — no server-side deployment, no API keys, no network requests after the initial model download. All inference runs on-device.

Key capabilities:
- **Prompt API** — General-purpose text generation with the `LanguageModel` interface
- **Multimodal input** — Image and audio understanding via the Prompt API
- **Summarizer API** — One-click text summarization
- **Writer API** — Generate text for a specific purpose
- **Rewriter API** — Rewrite existing text with different tone or length
- **Proofreader API** — Grammar and spelling correction
- **Language Detector API** — Detect the language of input text
- **Translator API** — Translate text between languages
- **Structured output** — JSON Schema-constrained responses
- **Session management** — Multi-turn conversations with context tracking
- **Chrome Extensions** — All APIs work in extensions

> [!NOTE]
> Gemini Nano is a **client-side** model. No data is sent to Google or any third party when using the model. The network is only required for the initial model download.

---

## Critical Rules (Always Apply)

### Hardware Requirements

The Prompt API, Summarizer, Writer, Rewriter, and Proofreader APIs require:
- **OS**: Windows 10/11, macOS 13+ (Ventura+), Linux, or ChromeOS (Chromebook Plus)
- **Storage**: At least 22 GB free on the volume containing the Chrome profile
- **GPU or CPU**: GPU with >4 GB VRAM, OR CPU with 16+ GB RAM and 4+ cores
- **Network**: Unmetered connection (only for initial model download)

Language Detector and Translator APIs work on Chrome desktop without the above GPU/RAM requirements.

### Supported Languages

From Chrome 140, Gemini Nano supports **English**, **Spanish**, and **Japanese** for input and output text.

### TypeScript Support

Install TypeScript definitions:
```bash
npm install @types/dom-chromium-ai
```

---

## Quick Start

### 1. Check Model Availability

Always check if the model is ready before creating a session:

```javascript
const availability = await LanguageModel.availability({

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The Chrome built-in AI APIs are accessed via the window.ai namespace (e.g., ai.languageModel, ai.summarizer). Using the capitalized interface names like LanguageModel as global entry points is incorrect and will result in a ReferenceError. This pattern should be updated throughout the document for all APIs (Summarizer, Writer, etc.).

Suggested change
const availability = await LanguageModel.availability({
const availability = await ai.languageModel.availability({
References
  1. Code examples in skill documentation should be minimal, demonstrating only the core SDK functionality. Avoid adding boilerplate like error handling, as the goal is to showcase SDK usage, not teach general coding practices.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is outdated and wrong now.

expectedInputs: [{ type: 'text', languages: ['en'] }],
expectedOutputs: [{ type: 'text', languages: ['en'] }],
});

// Returns: "unavailable" | "downloadable" | "downloading" | "available"
```

### 2. Create a Session

```javascript
const session = await LanguageModel.create({
monitor(m) {
m.addEventListener('downloadprogress', (e) => {
console.log(`Downloaded ${e.loaded * 100}%`);
});
},
});
```

> [!IMPORTANT]
> If the model isn't downloaded yet, the user must interact with the page first (click, tap, keypress) before `create()` can be triggered. Check `navigator.userActivation.isActive`.

### 3. Prompt the Model

#### Request-based (wait for full response)
```javascript
const result = await session.prompt('Explain quantum computing in simple terms.');
console.log(result);
```

#### Streaming (show partial results)
```javascript
const stream = session.promptStreaming('Write a poem about the ocean.');
for await (const chunk of stream) {
process.stdout.write(chunk);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

process.stdout.write is a Node.js-specific API and is not available in the browser or Chrome Extension environments where these APIs are used. Use console.log or a DOM-based output method instead.

Suggested change
process.stdout.write(chunk);
console.log(chunk);
References
  1. Code examples in skill documentation should be minimal, demonstrating only the core SDK functionality. Avoid adding boilerplate like error handling, as the goal is to showcase SDK usage, not teach general coding practices.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

+1, this should just be console.log(chunk);.

}
```

### 4. Stop a Prompt

```javascript
const controller = new AbortController();
stopButton.onclick = () => controller.abort();

const result = await session.prompt('Write a poem!', {
signal: controller.signal,
});
```

---

## Prompt API Features

### System Prompts and Context

Use `initialPrompts` to set system instructions and conversation history:

```javascript
const session = await LanguageModel.create({
initialPrompts: [
{ role: 'system', content: 'You are a helpful coding assistant.' },
{ role: 'user', content: 'What is the capital of France?' },
{ role: 'assistant', content: 'The capital of France is Paris.' },
],
});
```

### Response Prefix

Guide the model's response format with a prefix:

```javascript
const result = await session.prompt([
{ role: 'user', content: 'Create a JSON character sheet for a warrior' },
{ role: 'assistant', content: '```json\n', prefix: true },
]);
```

### Multimodal Input (Image + Audio)

```javascript
const session = await LanguageModel.create({
expectedInputs: [
{ type: 'text', languages: ['en'] },
{ type: 'image' },
{ type: 'audio' },
],
expectedOutputs: [{ type: 'text', languages: ['en'] }],
});

// Image input (supports Blob, HTMLImageElement, HTMLCanvasElement, etc.)
const imageBlob = await (await fetch('photo.jpg')).blob();
const result = await session.prompt([
{
role: 'user',
content: [
{ type: 'text', value: 'Describe what you see in this image:' },
{ type: 'image', value: imageBlob },
],
},
]);

// Audio input (supports AudioBuffer, ArrayBuffer, Blob, etc.)
const audioBuffer = await captureMicrophoneInput();
const audioResult = await session.prompt([
{
role: 'user',
content: [
{ type: 'text', value: 'Transcribe this audio:' },
{ type: 'audio', value: audioBuffer },
],
},
]);
```

> [!NOTE]
> Audio input requires a GPU. Supported input types: `AudioBuffer`, `ArrayBufferView`, `ArrayBuffer`, `Blob`. Image input supports: `HTMLImageElement`, `SVGImageElement`, `HTMLVideoElement`, `HTMLCanvasElement`, `ImageBitmap`, `OffscreenCanvas`, `VideoFrame`, `Blob`, `ImageData`.

### Structured Output (JSON Schema)

```javascript
const session = await LanguageModel.create();

const schema = {
type: 'object',
properties: {
sentiment: { type: 'string', enum: ['positive', 'negative', 'neutral'] },
confidence: { type: 'number' },
},
required: ['sentiment', 'confidence'],
};

const result = await session.prompt(
'Analyze the sentiment: "This product is amazing, I love it!"',
{ responseConstraint: schema }
);
console.log(JSON.parse(result));
// { sentiment: "positive", confidence: 0.95 }
```

### Append Messages

Pre-populate context without triggering inference:

```javascript
const session = await LanguageModel.create({
initialPrompts: [
{ role: 'system', content: 'You are a document analyst.' },
],
});

// Pre-load documents without generating a response
await session.append([
{
role: 'user',
content: [
{ type: 'text', value: 'Here is a document to analyze: ...' },
],
},
]);

// Now prompt with a question about the pre-loaded content
const analysis = await session.prompt('Summarize the key points.');
```

---

## Session Management

### Context Window Tracking

```javascript
console.log(`Context usage: ${session.contextUsage}/${session.contextWindow}`);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The LanguageModel session object uses tokensSoFar and maxTokens to track context usage, rather than contextUsage and contextWindow.

Suggested change
console.log(`Context usage: ${session.contextUsage}/${session.contextWindow}`);
console.log("Context usage: " + session.tokensSoFar + "/" + session.maxTokens);
References
  1. Code examples in skill documentation should be minimal, demonstrating only the core SDK functionality. Avoid adding boilerplate like error handling, as the goal is to showcase SDK usage, not teach general coding practices.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Outdated and wrong now.

```

### Context Overflow Handling

```javascript
session.addEventListener('contextoverflow', () => {
console.log('Context window exceeded — older messages will be dropped.');
});
```

### Clone a Session

```javascript
const clonedSession = await session.clone();
// Forked conversation preserves context and initial prompts
```

### Destroy a Session

```javascript
session.destroy();
// Frees resources. Session can no longer be used.
```

---

## Specialized APIs

### Summarizer API

```javascript
const summarizer = await Summarizer.create();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Specialized APIs should also be accessed via the ai namespace (e.g., ai.summarizer.create()) rather than using the interface name as a global.

Suggested change
const summarizer = await Summarizer.create();
const summarizer = await ai.summarizer.create();
References
  1. Code examples in skill documentation should be minimal, demonstrating only the core SDK functionality. Avoid adding boilerplate like error handling, as the goal is to showcase SDK usage, not teach general coding practices.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Outdated and wrong now.

const summary = await summarizer.summarize(longText);
```

### Writer API

```javascript
const writer = await Writer.create();
const text = await writer.write('Write a professional email declining a meeting.');
```

### Rewriter API

```javascript
const rewriter = await Rewriter.create();
const rewritten = await rewriter.rewrite('make this more formal: hey can u help me');
```

### Proofreader API

```javascript
const proofreader = await Proofreader.create();
const corrected = await proofreader.proofread('teh quikc brown fox jumpd');
```

### Language Detector API

```javascript
const detector = await LanguageDetector.create();
const result = await detector.detect('Bonjour le monde');
// { detectedLanguage: 'fr', confidence: 0.99 }
```

### Translator API

```javascript
const translator = await Translator.create({
sourceLanguage: 'en',
targetLanguage: 'es',
});
const translated = await translator.translate('Hello, how are you?');
// "Hola, ¿cómo estás?"
```

---

## Chrome Extensions

All built-in AI APIs work in Chrome Extensions. For extensions using the Prompt API, you can customize model parameters:

```javascript
const params = await LanguageModel.params();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The method to retrieve model parameters and limits is capabilities(), not params(). Additionally, it should be called on the ai.languageModel factory.

Suggested change
const params = await LanguageModel.params();
const params = await ai.languageModel.capabilities();
References
  1. Code examples in skill documentation should be minimal, demonstrating only the core SDK functionality. Avoid adding boilerplate like error handling, as the goal is to showcase SDK usage, not teach general coding practices.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Outdated and wrong now.

// { defaultTopK: 3, maxTopK: 128, defaultTemperature: 1, maxTemperature: 2 }

const session = await LanguageModel.create({
temperature: params.defaultTemperature,
topK: params.defaultTopK,
});
```

> [!NOTE]
> Remove the expired origin trial permission `"aiLanguageModelOriginTrial"` from your manifest if present.

---

## Enable on Localhost

1. Go to `chrome://flags/#optimization-guide-on-device-model` → **Enabled**
2. Go to `chrome://flags/#prompt-api-for-gemini-nano` → **Enabled** or **Enabled multilingual**
3. For multimodal input: `chrome://flags/#prompt-api-for-gemini-nano-multimodal-input` → **Enabled**
4. Restart Chrome
5. Verify: Open DevTools console, run `await LanguageModel.availability()`
Comment on lines +332 to +338

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Obsolete now that the API ships in Chrome 148.


---

## Best Practices

1. **Always check availability** before creating a session — the model may not be downloaded yet

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, and with the exact same options you'll pass to create().

2. **Inform users** about model download progress using the `monitor` callback
3. **Use streaming** for longer responses to provide immediate feedback
4. **Track context window usage** to handle overflow gracefully
5. **Destroy sessions** when no longer needed to free resources
6. **Use structured output** (JSON Schema) when you need parseable, predictable responses
7. **Prefer specialized APIs** (Summarizer, Writer, etc.) over the Prompt API for their specific tasks — they're optimized for those use cases
8. **Handle errors** — `QuotaExceededError` when context window is exceeded, `NotSupportedError` for unsupported modalities/languages
9. **Use `append()`** to pre-load context without triggering inference
10. **Test with real devices** — performance varies significantly across hardware

---

## Documentation Lookup

### Official Documentation

- [Get started with built-in AI](https://developer.chrome.com/docs/ai/get-started) — setup, requirements, model download
- [Prompt API](https://developer.chrome.com/docs/ai/prompt-api) — LanguageModel, multimodal, structured output
- [Session management](https://developer.chrome.com/docs/ai/session-management) — context window, cloning, destroying
- [Structured output](https://developer.chrome.com/docs/ai/structured-output-for-prompt-api) — JSON Schema constraints
- [Summarizer API](https://developer.chrome.com/docs/ai/summarizer-api) — text summarization
- [Writer API](https://developer.chrome.com/docs/ai/writer-api) — text generation
- [Rewriter API](https://developer.chrome.com/docs/ai/rewriter-api) — text rewriting
- [Proofreader API](https://developer.chrome.com/docs/ai/proofreader-api) — grammar correction
- [Language Detector API](https://developer.chrome.com/docs/ai/language-detection) — language detection
- [Translator API](https://developer.chrome.com/docs/ai/translator-api) — on-device translation
- [Cache models](https://developer.chrome.com/docs/ai/cache-models) — best practices for model caching
- [Debug Gemini Nano](https://developer.chrome.com/docs/ai/debug-gemini-nano) — troubleshooting
- [AI in Extensions](https://developer.chrome.com/docs/extensions/ai) — extension-specific guidance

### Demos

- [Prompt API Playground](https://chrome.dev/web-ai-demos/prompt-api-playground/)
- [Audio Prompt Demo](https://chrome.dev/web-ai-demos/mediarecorder-audio-prompt)
- [Image Prompt Demo](https://chrome.dev/web-ai-demos/canvas-image-prompt/)
- [Extension Sample](https://github.com/GoogleChrome/chrome-extensions-samples/tree/main/functional-samples/ai.gemini-on-device)