Cerebras, in an invisible overlay.

Cerebras inference runs on dinner-plate-sized chips and it shows — thousands of tokens per second on frontier open models. Connected to Pluely, even long answers land in a blink. Free tier available.

Your key, encrypted on-deviceUnmetered — free plan covers itOpenAI-compatible

Why Cerebras?

Wafer-scale speed, thousands of tokens per second.

The speed record holder

Cerebras regularly posts the fastest inference numbers in the industry — 2,000+ tokens per second on some models. Long summaries that take other providers fifteen seconds appear instantly.

Reasoning models you'll actually wait for

Thinking models are great until you're staring at a spinner mid-meeting. At Cerebras speeds, even chain-of-thought models respond in conversational time.

Free tier to feel the difference

The free tier is enough to run Pluely daily — try one meeting on it and normal-speed providers feel broken afterward.

Connection details

Base URL
https://api.cerebras.ai/v1
Example models
llama-4-maverickqwen3-32bgpt-oss-120b

Connect Cerebras in three steps

Any OpenAI-compatible endpoint works the same way.

01

Get your Cerebras API key

Create a key at cloud.cerebras.ai. It stays on your device — Pluely encrypts it locally and never sends it anywhere except Cerebras itself.

02

Add it as a custom provider

Open Pluely's dashboard → AI Providers → Add Custom Provider. Paste the base URL (https://api.cerebras.ai/v1) and your key, and pick the models you want available.

03

Pick it and ask

Your Cerebras models now appear in the overlay's model picker. Select one and every Ask, screenshot, and meeting answer runs on your key — unmetered by Pluely.

Everything Pluely does, on Cerebras

Connecting a provider doesn't limit the product — every mode routes through it.

Ask Cerebras about your screen

Capture the full screen, drag-select just the error or the clause that matters, attach PDFs and documents with built-in OCR, or turn on Use image so every message carries a fresh screenshot. Cerebras answers with the actual context in front of you.

Cerebras answers your meetings live

Listen mode transcribes your mic and system audio in real time with speaker labels, and Cerebras drafts the suggested answer before the question finishes — automatically on every question, after each pause, or only when you tap Suggest.

Summaries, transforms, and notes

One tap turns any answer, transcript, or document into a summary, key points with action items, a translation, or a longer draft — and meeting notes build themselves as you go. All of it runs on Cerebras.

Cerebras + Pluely — common questions

More providers that plug in

Put Cerebras over everything.

Install Pluely, paste your key once, and Cerebras answers for the next error, the next call, the next thing on your screen.

Free plan forever · Invisible on screen shares · No bot joins your calls

Explore Pluely

Download for your platform, browse release history, or explore our development journey

Platform Downloads

macOS

Apple Silicon & Intel

Download for macOS

Windows

x64 Architecture

Download for Windows

Linux

Debian Package

Download for Linux

Recent Releases

View All Releases

Loading releases...

Browse & Explore

All Downloads

Latest release downloads

View Downloads

Version History

Browse all releases

Browse Version History

Changelog

Development timeline

Browse Changelog

Ready to get started?

Download Pluely now and experience the privacy-first AI assistant that works seamlessly in the background.