Browser AI Playground · Try → Check → Build → Deploy

Try small open-weight AI models in your browser

One of our six free flagship tools. Load a small WebLLM model with WebGPU, test a prompt locally in your browser, then jump straight to the Compatibility Checker, PC Builder, model Wizard, or a stack recipe for your next step.

Start here

Where the Playground fits: Try → Check → Build

Start in the browser, then use the Checker and PC Builder to move from a quick test toward a real local setup.

What the Playground is for and how to read it

The Browser AI Playground lets you run a small open-weight model directly in your browser using WebGPU — no install, no account, and nothing sent to a server. It is for the curious first step: people who want to feel what local AI is like before committing to a full runtime such as Ollama or LM Studio. Read the results as a preview, not a performance test. The small models here load quickly but are far less capable than the larger ones you would run in a real local stack, so judge the workflow and the shape of the response rather than the raw answer quality. When a model feels promising, that is your cue to move to a full local setup where bigger models and longer context become possible.

Fast first test for most WebGPU laptops/desktops
Checking…

Model not loaded yet.

Checking browser support...

💬

No messages yet

Load the model, pick a starter prompt above, or type a question below. Treat answers as a lightweight demo, not verified guidance.

✓ Verified fallback available for this topic.

▸ Reference checklist (verified answers)
  • WebGPU and browser support: Use a current Chrome or Edge build first, confirm WebGPU is detected, and check the browser console if loading fails.
  • Model loading and speed: Expect the first model load to download hundreds of MB to several GB. Treat tokens/sec as a rough output-speed estimate, not a quality score.
  • Browser cache and storage: Browser caching of model files is expected. Clear site data if you need to reclaim storage or reset a failed download.
  • Prompt privacy limits: Prompts run in this browser session in the MVP. Normal hosting logs, analytics events, and third-party model download requests may still occur.
  • When to move to a full local stack: Use Ollama, LM Studio, or Open WebUI when you need larger models, stable model storage, RAG over files, repeatable APIs, multi-user use, or better GPU control.

What this playground does

  • Runs small model demos in the browser when WebGPU is available.
  • Shows model download progress, streaming output, and rough tokens/sec.
  • Keeps prompts client-side in this MVP.
  • Acts as a lightweight preview before following full local setup guides.

What this playground is good for

Good fit

  • Quick browser AI concept checks
  • Simple summaries and comparisons
  • Testing local browser inference feel
  • Deciding when to move to a full local stack

Use verified guides instead

  • Exact install commands
  • Hardware compatibility checks
  • Browser or GPU troubleshooting
  • Production workflow planning

Beta and privacy note: small browser models can be inaccurate, incomplete, slow, or unsupported on some devices. Check each model card and license before using outputs in production.

Included WebLLM model IDs

  • Llama 3.2 1B Instruct: Llama-3.2-1B-Instruct-q4f16_1-MLC · ~879MB VRAM required · 4k context
  • Llama 3.2 1B Instruct q4f32: Llama-3.2-1B-Instruct-q4f32_1-MLC · ~1.1GB VRAM required · 4k context
  • Llama 3.2 3B Instruct: Llama-3.2-3B-Instruct-q4f16_1-MLC · ~2.3GB VRAM required · 4k context
  • Llama 3.2 3B Instruct q4f32: Llama-3.2-3B-Instruct-q4f32_1-MLC · ~3GB VRAM required · 4k context
  • Llama 3.1 8B Instruct q4f16 1k: Llama-3.1-8B-Instruct-q4f16_1-MLC-1k · ~4.6GB VRAM required · 1k context

Continue: Try → Check → Build → Deploy

Move from browser testing to the right build path.

Use the Checker, PC Builder, model Wizard, and stack recipes when you need larger models, longer context, or repeatable local setup.