Wiring Page Assist to Ollama and a Cloud Gateway: Setup, Four Workflows, and the Gotchas
TL;DR : Page Assist is an MIT-licensed browser extension that puts a model in a sidebar next to the page you're reading. Ollama on localhost:11434 is detected automatically. Cloud providers are added from a preset dropdown; since v1.5.86 that includes AIHubMix, so connecting it is just picking it and pasting an API key. Short, private tasks run fine on a local model. Long pages and multi-step…
Page Assist is an MIT-licensed browser extension that integrates a model into a sidebar next to the page you're reading. The extension automatically detects Olloma running on localhost:11434. For cloud providers, you can select from a preset dropdown; starting from version 1.5.86, AIHubMix is one of the available options. Connecting AIHubMix involves obtaining an API key and entering it in the settings.
Page Assist streamlines the process of using a model by collapsing the typical workflow of copy-pasting text, pasting the answer back, into a sidebar and a full-tab web UI. The extension functions as a client, calling the Chat Completions endpoint to chat and storing history, settings, and knowledge-base embeddings in browser storage.
It does not send any telemetry. As of October 4, 2026, Page Assist has over 8,200 GitHub stars and 300,000 users on the Chrome Web Store. The prerequisites for using Page Assist include a Chromium browser or Firefox, with Opera and Arc supporting only the web UI. Optional components include Olloma for local models and an OpenAI-compatible cloud endpoint.
The installation process involves starting Olloma on the default port and opening Page Assist, where pulled models appear without any configuration. If a 403 error occurs on sending, it may be due to CORS. Two potential fixes include setting the OLLAMA_ORIGINS environment variable or enabling the option to configure custom origin URLs under the Settings.
To add AIHubMix from the preset list, ensure you have the API key and then enter it into the OpenAI-compatible provider settings. In the model list, select the desired chat model and save the changes. For the embedding model, you can use a local nomic-embed-text model via Olloma or choose a cloud embedding model like gemini-embedding-001 from AIHubMix.
It's important to note that the selected model should be an embedding model, not a chat model. The extension offers four workflows: 1. Chat with Website, which has two modes: embedding and retrieval. The embedding mode chunks the page text and runs it through an embedding model, sending the top chunks for retrieval. The full context mode sends the entire page text directly.
The latter option is recommended for better summaries and cross-section questions, while it requires a model with a larger context window. Additional features include @tab mentions, which allow pulling content from other tabs into a single prompt, a YouTube Summarize button, and Vision mode for summarizing images. 2. Copilot right-click prompts include built-in actions such as Summarize, Rephrase, Translate, Explain, and Custom.
The Custom Copilot Prompts can be created by providing a title and a template with a {text} placeholder, resulting in separate context-menu entries. These prompts are suitable for local models due to their short and frequent use, eliminating per-call costs and network hops. 3. Knowledge Base allows users to add new knowledge sources in various formats such as .pdf, .docx, .txt, .csv, and .md.
The processing and vector storage occur in the browser, which may cause performance issues with large collections. If retrieval fails to find relevant passages, upgrading the embedding model before using the chat model is recommended. 4. Page Action and MCP (Multi-Model Collaboration) enable remote execution of the model on separate servers.
The Page Action is a Chromium-only companion extension that requires the debugger permission. It reads the tab, then performs actions such as clicking, typing, scrolling, and filling forms step by step, with approval required before each action by default. The MCP tools run without approval for servers that can only read or write data.
Multi-step tool use, which can lead to stalls or loops, should utilize a capable cloud model. Failure modes include a 403 error on sending, which is typically due to CORS. To resolve this, set the OLLAMA_ORIGINS environment variable or enable custom origin URLs in the Settings.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.