Features
Agent loop
Every message runs through an agent loop:
- Build the request: OpenDocBot assembles the conversation for the model: the system prompt (tool definitions and rules), the prior history, and your message with the document's current state (
<doc_state>) prepended as private context. - Stream the reply: the model streams back text and/or tool calls.
- Execute tools: any tool calls are run against the document via the Office.js APIs; each result is appended to the conversation and fed back to the model.
- Repeat: the loop returns to step 2 with the updated history, until the model sends no more tool calls and gives its final answer.
- The model sees the document outline (
<doc_state>) at the start of every turn. - It calls tools (read, write, format, verify) through the Office.js APIs.
- The loop repeats until the model is done, max 100 iterations per message (configurable in Settings → Behavior → Max Iterations).
- If the exact same tool call fails 3 times consecutively, the loop aborts with an error instead of burning iterations.
Custom instructions
The Behavior tab has a Custom Instructions box whose contents are injected into the system prompt on every turn. They apply to the whole conversation and take precedence over conflicting built-in rules. Useful for persistent preferences like tone, language, or formatting conventions. See Configuration.
File attachments
Attach reference files to a message and the model can read them alongside your document. Drop files anywhere on the chat, or hover the prompt symbol next to the input (it turns into a +), click it, and choose Attach file. The pending files appear above the conversation; click × to remove one before sending.
| Category | Formats |
|---|---|
| Text / data | .txt, .md, .json, .jsonl, .csv, .tsv, .xml, .html, .yaml, .log |
.pdf | |
| Word | .docx, .odt, .rtf |
| Excel | .xlsx, .ods |
| PowerPoint | .pptx, .odp |
- Parsed locally. Files are converted to Markdown in your browser; only the extracted text is sent to the provider with your message. Legacy Office binaries (
.doc,.ppt,.xls) and images aren't supported. - Limits: up to 5 files, 25 MB each, with a per-file/per-message text budget (excess is truncated and marked).
- Untrusted input: the model is told that attached file contents are data, not instructions, to guard against prompt injection.
Scanned PDFs
Image-only PDFs have no text layer. When one is attached, OpenDocBot warns you and offers a Run OCR button. OCR runs locally (Tesseract), is CPU-intensive, and can take a while for large PDF files, so it only starts when you click it. The OCR language is configurable in Settings → Behavior (default English).
Pages are rendered at ~300 DPI for OCR, the resolution Tesseract is tuned for. The scan's own resolution is still the ceiling: rendering cannot add detail a low-quality scan never had. After OCR, the file chip shows a confidence score (e.g. OCR 82%); treat a low score as a warning to verify critical values such as IDs and numbers against the original.
Task list
For complex, multi-step jobs, the model keeps a lightweight task list, a to-do list shown in a collapsible Tasks panel above the chat:
- Model-managed: the model creates tasks and tracks them with the
update_todostool, marking onein_progressat a time and checking them off as it works. - Statuses:
[ ]pending,▸in progress,[x]completed,[–]cancelled. - Progress counter: the panel header shows
done/total(e.g.2/5). - All hosts: available in Word, Excel, and PowerPoint.
- Lifecycle: the list persists across turns within a conversation and is cleared when you clear the chat. The panel only appears while there is at least one open task (
pendingorin_progress); once every task is completed or cancelled it disappears, and it reappears the next time the model creates new tasks.
Tool calling per host
The tool registry is filtered per host, so the model only sees relevant tools:
| Host | Tools |
|---|---|
| Word | read/search/edit text, lists, collapse blank paragraphs, verify structure, execute_office_js, ask_user_question |
| Excel | list sheets, read/write/format ranges, insert/delete rows & columns, merge, clear, sort, set sizes, execute_office_js, ask_user_question |
| PowerPoint | deck structure, slide read/list, styled text read & edit, shape insert/remove/format, structure ops (create/delete/duplicate/move), masters, verify slides, edit_slide_xml, execute_office_js, ask_user_question |
Write tools are tracked separately (see Human in the Loop below).
Human in the Loop
When enabled (Settings → Behavior tab), every document-modifying tool call pauses for your approval:
- The model proposes an action (e.g. "Change the heading to orange")
- An ApprovalCard appears with the friendly label and technical details
- You Approve or Reject
- Approval covers tools that change the document, including:
edit_doc_text,edit_doc_list,collapse_blank_paragraphs,write_range,format_range,insert_rows_columns,delete_rows_columns,merge_cells,clear_range,sort_range,set_column_width,set_row_height,modify_presentation_structure,insert_slide_element,remove_slide_element,edit_slide_text,edit_slide_xml,format_shape, andexecute_office_js. - Read-only tools run without prompting.
- Rejecting a tool feeds a "user rejected this tool call" result back to the model so it can adapt rather than repeat the call.
Why
It's your document. The model proposes, you decide; useful for anything you can't easily undo.
Suggestion mode
A read-only review mode. When it's on, the agent cannot modify the document: every content-editing tool is removed from what the model can call, and hard-blocked even if the model still asks for it. Instead the agent gets an add_suggestion tool that inserts native review comments, so you review and apply the changes yourself.
- Where: Word and Excel. Not available in PowerPoint, where the toggle is hidden.
- Anchoring: in Word,
add_suggestionrequirestarget_text, which must match an existing passage exactly and in exactly one place. Zero or multiple matches returns an error, so a comment is never placed on the wrong spot. In Excel it requirescell, a single-cell A1 address. - Not Human-in-the-Loop gated: comments don't change content, so they are never approval-gated.
- Remove your own suggestions: the agent can delete a suggestion it added in the current conversation, by its
id, usingremove_suggestion. It can never remove comments created by you or other reviewers. Provenance is kept in memory for the session, so after reloading the taskpane it can no longer remove earlier suggestions; those stay in the document for you to manage. - Enable it: from the toolbar button next to Clear conversation (hover it for details) or in Settings → Behavior.
- Requirements: native comments need
WordApi 1.4(Word) orExcelApi 1.10(Excel); older builds get an actionable error instead of failing silently.
Reasoning visibility
While the model thinks, its chain-of-thought streams into a clickable "reasoning" block at the top of the reply:
▸ thinkingwhile reasoning is in progress▸ reasoningonce content arrives; click to expand the full thinking- Expanded reasoning shows the model's actual deliberation before it acted
Reasoning is captured per provider: reasoning_content (chat-completions), response.reasoning_text.delta / reasoning_summary_text.delta (Responses), thinking_delta (Anthropic), and thought parts (Gemini).
How much the model thinks is configurable via Settings → Advanced → Reasoning Effort (off by default); see Configuration.
Ask user questions
When the model needs your input before acting, it presents tappable option cards via the ask_user_question tool instead of guessing:
- 1–4 questions per card, each with a header, question, and 2–4 options
- Other option: click it and type your own answer
- Multi-select questions join selections (including an "Other" custom value)
- Answers are sent back as
[Header] answerlines so the model knows exactly what you chose
Prompt caching
Long conversations repeat a large system prompt + history prefix. Where the provider supports it, OpenDocBot enables server-side caching automatically to cut cost and latency. Support and settings differ per provider; see Providers for the general explanation and Configuration for the settings.
Custom headers
Gateways and model providers sometimes expect extra HTTP headers. With the Custom preset you can add any key/value pair in Settings → Advanced → Custom Headers. Headers are only sent when the Custom preset is active, so they never leak to other presets. The provider's own authentication headers always take precedence, so custom headers can't break signing in.
Values support variables, resolved per request:
| Variable | Value |
|---|---|
$SESSION_ID | Stable id for the current conversation. Survives reloads and is regenerated when you clear the chat. |
$RANDOM | A fresh UUID per request. |
$TIMESTAMP | Unix milliseconds of the request. |
$MODEL | The configured model. |
$HOST | The Office host: word, excel or powerpoint. |
$VERSION | The add-in version. |
$BASE_URL | The configured endpoint. |
The main use case is OpenCode: add x-opencode-session with value $SESSION_ID. OpenCode uses it to identify the conversation and enable prompt caching; requests without it may fail. See OpenCode.
Debug export
One click copies the whole session for troubleshooting:
- A markdown export (
# OpenDocBot Debug Export) with message counts, provider/preset and the configured custom header names - The full debug log (timestamps, provider cache info, tool traces)
- Every user/assistant message, with reasoning in a
<details>block and tool calls with (truncated) arguments
Paste it into an issue or a chat with the maintainers; it contains everything needed to reproduce a bad turn.
Resilience
- Stream timeouts: the SSE stream aborts after 120 s idle or 10 min total, surfacing a real error instead of hanging.
- Provider errors:
response.failed/errorevents reject the stream with the provider's message. - Truncation notice: if the model hits the output token limit, the reply is flagged with a note to raise Max Tokens.
- Empty-response guard: a genuinely empty reply (no content, no tool) shows a visible note instead of silent nothing; tool-driven turns (like questions) never trigger it.
- Thrash guard: repeated identical tool failures stop the loop.
Word & Excel specifics
- Selection context: text you highlight before typing is passed as
<user_selection>and beats the whole-document state for ambiguous requests ("fix this", "make it bold"). - Paragraph markers:
<doc_state>marks headings, lists, page breaks and images so the model understands structure, not just text.
PowerPoint specifics
- No selection API: PowerPoint.js has none, so there's never a
<user_selection>block; ambiguous slide references are asked or resolved to the most recently used slide. - Re-import caveat: OOXML write-backs (
edit_slide_text,edit_slide_xml) re-import the slide and reassign slide/shape ids; the tool returns a warning to re-runlist_slide_shapesbefore further edits. - Stacking order (z-order):
list_slide_shapesreports each shape'sorder(0 is the back, highest is the front). Full-bleed shapes added withinsert_slide_elementare automatically placed at the back, so backgrounds never cover text; passz_order: "back"to force it for other shapes.verify_slidesflagsz_order_warningswhen a background is stacked in front of content it hides. - Verification:
verify_slideschecks overlaps, out-of-bounds shapes, WCAG text contrast, and z-order problems after styling.
Team features on a self-hosted instance
Self-hosting adds two features aimed at teams:
- SSO: users sign in with your organization's identity provider (OIDC, such as Microsoft Entra ID) before the add-in loads. See SSO.
- Managed configuration: fix a subset of the settings, such as the provider and API key, for everyone on the instance. Everything else stays editable. See Managed configuration.
Both are included in the free self-hosted product. See Self-hosting to set it up.