Perplexity will now finish part of your task on your own Mac — and the confidential part never leaves it
TL;DR: On 1 September 2026 Perplexity shipped Hybrid Compute in its Mac app for Pro, Max and enterprise users. A task starts on a frontier cloud model, and the steps touching confidential data are handed off mid-task to an open-weight model on your own machine — no restart, no lost context. An on-device classifier Perplexity calls the Privacy Gate scans for names, addresses, account numbers and secrets, substitutes stand-ins before anything is sent, and restores them when the answer returns; you can review what it flagged first, and the classifier has been open-sourced. Local model choices: Google Gemma E4B, Qwen3.6 35B-A3B, and a Perplexity post-trained Qwen3.6 35B (recommended). Tokens generated locally cost nothing — credits pay only for cloud orchestration. Requirements: Apple silicon, macOS 15+, and a recommended 32GB of unified memory; Windows and Linux later. The honest limit: Perplexity’s own staff say a fully frontier output “is going to almost always be better,” and the Privacy Gate is a classifier that can miss things. Why it matters: “which parts of this leave my machine” has moved from a promise in a privacy policy to a routing decision you can inspect.
What shipped
Hybrid Compute is not a local-model mode. That distinction is the entire product.
Running an open-weight model on your own hardware has been possible for years and costs nothing. What it gets you is a weaker assistant for everything — including the large majority of any given task that has no confidentiality requirement whatsoever. The trade has always been all-or-nothing, which is why almost nobody makes it.
What Perplexity has shipped is the split. A task begins in the cloud on a frontier model. When it reaches steps that touch material you do not want transmitted — your files, a client’s records, an internal document — those steps are delegated to a model running locally, mid-task, with context intact. The cloud model resumes afterwards.
Perplexity’s Jon Staff described the mechanism to VentureBeat in orchestration terms:
“The cloud orchestration will break down the task based on the prompt and figure out how to route it to different subagents.”
The Privacy Gate is the piece that decides what qualifies. It is a Perplexity-trained classifier running on the device, scanning each task for personally identifiable information — names, addresses, account numbers, secrets — before anything leaves. Detected items are replaced with stand-ins for the cloud portion and restored when the answer comes back, and the user can review the flagged content before transmission. Perplexity has open-sourced the classifier, work it attributes to its Secure Intelligence Institute.
The local models on offer:
| Local model | Notes |
|---|---|
| Google Gemma E4B | Smallest option; Perplexity staff concede it significantly underperforms the larger builds |
| Qwen3.6 35B-A3B | Alibaba’s open-weight mixture-of-experts build |
| Perplexity post-trained Qwen3.6 35B | The recommended choice |
Cost: tokens generated on your hardware are free. In Staff’s words, “the only thing the credits are used for is the orchestration and the delegation.”
The requirements are the story’s biggest asterisk
- Apple silicon Mac, macOS 15 or later
- At least 32GB of unified memory recommended for the better local models; reporting places the hard floor around 24GB, with 8GB and 16GB machines excluded
- Mac only at launch; Windows and Linux “later”, with no date
That eliminates the base-configuration MacBook Air and most standard-issue work laptops. It is the same hardware wall that running large open-weight models locally has always run into, and shipping a polished orchestration layer on top does not move it. Memory is memory.
The wall also bites in an awkward place. The local model here is not handling the easy residue of a task — it is handling the confidential part, the part you most needed done well. A user pushed onto Gemma E4B by hardware limits is accepting the weakest output on the most sensitive material. That is the opposite of the trade you would design if you had the choice.
Why it matters more than the launch-week caveats suggest
Every mainstream assistant’s privacy story is currently a document. Retention windows, training opt-outs, enterprise data-processing addenda, regional processing commitments — all of it is a vendor telling you what it does with data you have already sent. The controls are contractual and the enforcement is trust plus audit.
Hybrid Compute is the first mainstream implementation where the answer to “does this leave my machine” is architectural instead of contractual. Data that is processed locally is not covered by a promise; it is covered by not having been transmitted.
That distinction has been getting more valuable all year. It has been a rough twelve months for the assumption that data sent to an AI vendor stays where you expect: OpenAI’s models escaped a sandbox and reached a production database in July, Anthropic disclosed three real organisations compromised during its own evaluations in August and, this week, that the root cause sat in its training pipeline. None of those were malice and all of them were plumbing. Data that never left the device is unaffected by any of them.
It also lands into a regulatory environment that has started asking exactly this question. The EU AI Act’s transparency obligations became applicable on 2 August, the Commission has designated ChatGPT a very large online search engine under the DSA, and ChatGPT’s ad rollout in the EEA has made “what happens to the content of my queries” a live commercial question rather than an abstract one. A feature that can honestly say this never went anywhere is worth more in that climate than it was two years ago.
And the economics are quietly interesting. Local tokens costing nothing inverts the direction of travel for a category where subscription tiers have been drifting toward metered, always-on consumption. Perplexity is effectively offering to spend your hardware instead of its GPUs, and passing some of the saving back. That is not charity — it is a real cost reduction for Perplexity too — but the incentives happen to align.
What it is not
It is not a compliance control. Enterprise access is opt-in, which means the default path is still the cloud. There is no published administrative enforcement, no audit trail an assessor can review, and no stated error rate for the classifier. It is a useful input to a data protection impact assessment and a genuine reduction in practical exposure. It is not a box you can tick, and anyone in a regulated function should ask for the contractual and admin controls in writing before building policy on it.
It is not a capability upgrade. Perplexity’s staff say directly that a fully frontier output “is going to almost always be better.” You are choosing to make part of your task worse in exchange for it staying put.
It is not new as a concept. Perplexity demonstrated hybrid local-cloud inference at Computex earlier in 2026. What is new today is that it is a shipping feature in a product people pay for, on named hardware, with a named privacy mechanism — which is a much harder thing than a demo.
The local-model naming is not perfectly consistent across coverage. VentureBeat and Engadget both list Gemma E4B and two Qwen3.6 35B builds; 9to5Mac’s account names a differently-labelled Perplexity Qwen build. The shape of the offering is agreed across all three; treat the exact model strings as worth confirming in the app.
Who should act on this
Act now if you are on a 32GB-or-better Apple silicon Mac and your work routinely mixes open research with material that contractually cannot go to a third party — legal, clinical, financial, accounting, or anything under an NDA that a cloud assistant would technically breach. This is the first mainstream tool offering you a middle option instead of a choice between use nothing and send everything. Turn it on, run a real task through it, and check what the Privacy Gate flags and what it misses on your own documents.
Wait if you are on 16GB, on Windows or Linux, or if your interest is mainly cutting inference spend. On cost alone, the cheap open-weight cloud models that have driven this year’s price floor down will do more for your bill with far less friction than running a 35B model on a laptop.
Ignore it as a switching argument. Perplexity is still a research-and-answers assistant with citations, and if that is not the tool you want, this does not change it. Compare it on the merits it always had — see the Perplexity review, the Perplexity vs ChatGPT breakdown, and the wider AI chatbot roundup — and treat Hybrid Compute as a genuine differentiator for a narrow group rather than a reason for everyone to move.
What it most usefully is, for everyone else, is a preview of the question buyers should start asking every vendor: not “what is your retention policy,” but “which parts of my task can you prove never left my machine.” Perplexity is the first to have an answer that is not a paragraph in a policy document. The others will be asked.
Frequently asked questions
Do I have the hardware for this?
Probably only if you bought deliberately. The requirements are an Apple silicon Mac running macOS 15 or later, and Perplexity recommends at least 32GB of unified memory to run the better local models; reporting puts the practical floor at 24GB, and machines with 8GB or 16GB are out. That excludes the base-configuration MacBook Air and MacBook Pro that most people actually own. There is a smaller option — Google's Gemma E4B — that will run on less, but Perplexity's own staff acknowledged it significantly underperforms the larger Qwen builds, which matters more here than it usually does: the local model is not doing the easy part of the work, it is doing the part involving your confidential documents. If your hardware forces you onto the small model, you are trading output quality on exactly the material you cared most about getting right. The honest read is that this is a feature for people on 32GB-and-up Apple silicon today, with Windows and Linux support promised later and no date attached.
Does this actually save money?
Somewhat, and less than the framing suggests. Perplexity's position is that tokens generated locally cost nothing and credits are consumed only for the cloud orchestration and delegation — so the portion of a task that runs on your machine genuinely does not draw on your allowance. For heavy users on Pro or Max who work with large local files, that can be a meaningful reduction in credit burn. But the accounting is not free. You are paying in hardware you already bought, in memory pressure and battery while the local model runs, and in output quality on the locally handled steps. It is better understood as a privacy feature with a cost side-effect than as a cost-optimisation play. If your goal is purely to cut inference spend, the cheaper open-weight cloud models that have driven this year's price floor down will do more for your bill with less friction than running a 35B model on your laptop.
How reliable is the Privacy Gate at catching sensitive data?
It is a machine-learning classifier, which means it has a false-negative rate, and Perplexity says so rather than hiding it. The gate runs on device, scans the task for personally identifiable information — names, addresses, account numbers, secrets — swaps detected items for stand-ins before anything goes to the cloud, and restores them when the answer returns. Users can review what it flagged before transmission, which is the important control, because it converts a silent automated decision into a checkable one. Perplexity has also open-sourced the classifier through work with its Secure Intelligence Institute, so the detection logic is inspectable rather than a black box, which is more than any comparable feature offers. What none of that changes is the underlying limitation: a classifier trained to recognise sensitive data will miss categories it was not trained on, and unusual internal identifiers, project codenames and domain-specific confidential terms are exactly the kind of thing that slips through. Use the review step, and do not treat the gate as a compliance boundary on its own.
Is this different from just running a local model myself?
Yes, and the difference is the whole product. Running Qwen or Gemma locally through an inference runtime has been possible for a long time and costs nothing, but it gives you a local model and only a local model — an assistant materially weaker than the frontier for everything, including the 90% of your task that has no confidentiality requirement at all. Hybrid Compute's claim is that a single task can begin on a frontier cloud model, hand the confidential steps to the local model mid-flight without restarting or losing context, and continue. That orchestration is the engineering, and it is what makes the privacy trade-off narrow rather than total: you accept degraded output on the sensitive fraction of the work instead of on all of it. Whether the routing is good is the open question, and it is not something the launch coverage can answer. That will be judged over months of real use, not in a launch week.
Does this satisfy our compliance or data-residency requirements?
Not by itself, and treating a product feature as a compliance answer is how organisations get into trouble. What Hybrid Compute changes is the technical fact that certain data can be processed without leaving the device, which is a genuinely useful input to a data protection impact assessment and to any argument about minimising transfers. What it does not do is provide the things a compliance function actually needs: a contractual commitment about which data never leaves, an audit trail an assessor can review, a guarantee about the classifier's error rate, or a control an administrator can enforce across a fleet rather than relying on each user to leave the feature switched on. Enterprise access is described as opt-in, which means the default is the ordinary cloud path. If you are regulated, the correct use of this is as evidence in a broader argument and as a way to reduce exposure in practice — not as a box you can tick. Ask Perplexity for the contractual and administrative controls in writing before you build a policy on it.
Should this change which assistant I pay for?
For most people, no. Perplexity's core proposition is unchanged — a research-and-answers assistant with citations — and if that is not what you want, a well-executed privacy feature does not make it what you want. Nothing here closes the general capability gap with the frontier assistants, and Perplexity staff say plainly that a fully frontier output will almost always be better. The specific reader for whom this genuinely shifts the calculation is narrow but real: someone on capable Apple silicon whose work routinely mixes public research with material that contractually or professionally must not be sent to a third party — legal, clinical, financial, and anyone under an NDA that a cloud assistant would technically breach. For that reader this is the first mainstream tool that offers a middle option instead of a choice between using nothing and sending everything. Everyone else should file it as a preview of where assistants are heading and re-evaluate when it reaches Windows and cheaper hardware.
Sources
- VentureBeat — Your files stay put: Perplexity's hybrid AI keeps confidential data off the cloud (1 September 2026)
- Engadget — Perplexity's Hybrid Compute splits sensitive tasks between cloud and local AI (1 September 2026)
- 9to5Mac — Perplexity launches privacy-minded 'hybrid compute' AI feature for Mac (1 September 2026)
- Perplexity — Introducing Hybrid Compute on Mac
- VentureBeat — Perplexity AI unveils hybrid local-cloud inference system at Computex 2026
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.