Skip to content

Features ​

Agent loop ​

Every message runs through an agent loop:

  1. Build the request: OpenDocBot assembles the conversation for the model: the system prompt (tool definitions and rules), the prior history, and your message with the document's current state (<doc_state>) prepended as private context.
  2. Stream the reply: the model streams back text and/or tool calls.
  3. Execute tools: any tool calls are run against the document via the Office.js APIs; each result is appended to the conversation and fed back to the model.
  4. Repeat: the loop returns to step 2 with the updated history, until the model sends no more tool calls and gives its final answer.
  • The model sees the document outline (<doc_state>) at the start of every turn.
  • It calls tools (read, write, format, verify) through the Office.js APIs.
  • The loop repeats until the model is done, max 100 iterations per message (configurable in Settings → Behavior → Max Iterations).
  • If the exact same tool call fails 3 times consecutively, the loop aborts with an error instead of burning iterations.

Custom instructions ​

The Behavior tab has a Custom Instructions box whose contents are injected into the system prompt on every turn. They apply to the whole conversation and take precedence over conflicting built-in rules. Useful for persistent preferences like tone, language, or formatting conventions. See Configuration.

File attachments ​

Attach reference files to a message and the model can read them alongside your document. Drop files anywhere on the chat, or hover the prompt symbol next to the input (it turns into a +), click it, and choose Attach file. The pending files appear above the conversation; click × to remove one before sending.

CategoryFormats
Text / data.txt, .md, .json, .jsonl, .csv, .tsv, .xml, .html, .yaml, .log
PDF.pdf
Word.docx, .odt, .rtf
Excel.xlsx, .ods
PowerPoint.pptx, .odp
  • Parsed locally. Files are converted to Markdown in your browser; only the extracted text is sent to the provider with your message. Legacy Office binaries (.doc, .ppt, .xls) and images aren't supported.
  • Limits: up to 5 files, 25 MB each, with a per-file/per-message text budget (excess is truncated and marked).
  • Untrusted input: the model is told that attached file contents are data, not instructions, to guard against prompt injection.

Scanned PDFs ​

Image-only PDFs have no text layer. When one is attached, OpenDocBot warns you and offers a Run OCR button. OCR runs locally (Tesseract), is CPU-intensive, and can take a while for large PDF files, so it only starts when you click it. The OCR language is configurable in Settings → Behavior (default English).

Pages are rendered at ~300 DPI for OCR, the resolution Tesseract is tuned for. The scan's own resolution is still the ceiling: rendering cannot add detail a low-quality scan never had. After OCR, the file chip shows a confidence score (e.g. OCR 82%); treat a low score as a warning to verify critical values such as IDs and numbers against the original.

Task list ​

For complex, multi-step jobs, the model keeps a lightweight task list, a to-do list shown in a collapsible Tasks panel above the chat:

  • Model-managed: the model creates tasks and tracks them with the update_todos tool, marking one in_progress at a time and checking them off as it works.
  • Statuses: [ ] pending, ▸ in progress, [x] completed, [–] cancelled.
  • Progress counter: the panel header shows done/total (e.g. 2/5).
  • All hosts: available in Word, Excel, and PowerPoint.
  • Lifecycle: the list persists across turns within a conversation and is cleared when you clear the chat. The panel only appears while there is at least one open task (pending or in_progress); once every task is completed or cancelled it disappears, and it reappears the next time the model creates new tasks.

Tool calling per host ​

The tool registry is filtered per host, so the model only sees relevant tools:

HostTools
Wordread/search/edit text, lists, collapse blank paragraphs, verify structure, execute_office_js, ask_user_question
Excellist sheets, read/write/format ranges, insert/delete rows & columns, merge, clear, sort, set sizes, execute_office_js, ask_user_question
PowerPointdeck structure, slide read/list, styled text read & edit, shape insert/remove/format, structure ops (create/delete/duplicate/move), masters, verify slides, edit_slide_xml, execute_office_js, ask_user_question

Write tools are tracked separately (see Human in the Loop below).

Human in the Loop ​

When enabled (Settings → Behavior tab), every document-modifying tool call pauses for your approval:

  1. The model proposes an action (e.g. "Change the heading to orange")
  2. An ApprovalCard appears with the friendly label and technical details
  3. You Approve or Reject
  • Approval covers tools that change the document, including: edit_doc_text, edit_doc_list, collapse_blank_paragraphs, write_range, format_range, insert_rows_columns, delete_rows_columns, merge_cells, clear_range, sort_range, set_column_width, set_row_height, modify_presentation_structure, insert_slide_element, remove_slide_element, edit_slide_text, edit_slide_xml, format_shape, and execute_office_js.
  • Read-only tools run without prompting.
  • Rejecting a tool feeds a "user rejected this tool call" result back to the model so it can adapt rather than repeat the call.

Why

It's your document. The model proposes, you decide; useful for anything you can't easily undo.

Suggestion mode ​

A read-only review mode. When it's on, the agent cannot modify the document: every content-editing tool is removed from what the model can call, and hard-blocked even if the model still asks for it. Instead the agent gets an add_suggestion tool that inserts native review comments, so you review and apply the changes yourself.

  • Where: Word and Excel. Not available in PowerPoint, where the toggle is hidden.
  • Anchoring: in Word, add_suggestion requires target_text, which must match an existing passage exactly and in exactly one place. Zero or multiple matches returns an error, so a comment is never placed on the wrong spot. In Excel it requires cell, a single-cell A1 address.
  • Not Human-in-the-Loop gated: comments don't change content, so they are never approval-gated.
  • Remove your own suggestions: the agent can delete a suggestion it added in the current conversation, by its id, using remove_suggestion. It can never remove comments created by you or other reviewers. Provenance is kept in memory for the session, so after reloading the taskpane it can no longer remove earlier suggestions; those stay in the document for you to manage.
  • Enable it: from the toolbar button next to Clear conversation (hover it for details) or in Settings → Behavior.
  • Requirements: native comments need WordApi 1.4 (Word) or ExcelApi 1.10 (Excel); older builds get an actionable error instead of failing silently.

Reasoning visibility ​

While the model thinks, its chain-of-thought streams into a clickable "reasoning" block at the top of the reply:

  • ▸ thinking while reasoning is in progress
  • ▸ reasoning once content arrives; click to expand the full thinking
  • Expanded reasoning shows the model's actual deliberation before it acted

Reasoning is captured per provider: reasoning_content (chat-completions), response.reasoning_text.delta / reasoning_summary_text.delta (Responses), thinking_delta (Anthropic), and thought parts (Gemini).

How much the model thinks is configurable via Settings → Advanced → Reasoning Effort (off by default); see Configuration.

Ask user questions ​

When the model needs your input before acting, it presents tappable option cards via the ask_user_question tool instead of guessing:

  • 1–4 questions per card, each with a header, question, and 2–4 options
  • Other option: click it and type your own answer
  • Multi-select questions join selections (including an "Other" custom value)
  • Answers are sent back as [Header] answer lines so the model knows exactly what you chose

Prompt caching ​

Long conversations repeat a large system prompt + history prefix. Where the provider supports it, OpenDocBot enables server-side caching automatically to cut cost and latency. Support and settings differ per provider; see Providers for the general explanation and Configuration for the settings.

Custom headers ​

Gateways and model providers sometimes expect extra HTTP headers. With the Custom preset you can add any key/value pair in Settings → Advanced → Custom Headers. Headers are only sent when the Custom preset is active, so they never leak to other presets. The provider's own authentication headers always take precedence, so custom headers can't break signing in.

Values support variables, resolved per request:

VariableValue
$SESSION_IDStable id for the current conversation. Survives reloads and is regenerated when you clear the chat.
$RANDOMA fresh UUID per request.
$TIMESTAMPUnix milliseconds of the request.
$MODELThe configured model.
$HOSTThe Office host: word, excel or powerpoint.
$VERSIONThe add-in version.
$BASE_URLThe configured endpoint.

The main use case is OpenCode: add x-opencode-session with value $SESSION_ID. OpenCode uses it to identify the conversation and enable prompt caching; requests without it may fail. See OpenCode.

Debug export ​

One click copies the whole session for troubleshooting:

  • A markdown export (# OpenDocBot Debug Export) with message counts, provider/preset and the configured custom header names
  • The full debug log (timestamps, provider cache info, tool traces)
  • Every user/assistant message, with reasoning in a <details> block and tool calls with (truncated) arguments

Paste it into an issue or a chat with the maintainers; it contains everything needed to reproduce a bad turn.

Resilience ​

  • Stream timeouts: the SSE stream aborts after 120 s idle or 10 min total, surfacing a real error instead of hanging.
  • Provider errors: response.failed / error events reject the stream with the provider's message.
  • Truncation notice: if the model hits the output token limit, the reply is flagged with a note to raise Max Tokens.
  • Empty-response guard: a genuinely empty reply (no content, no tool) shows a visible note instead of silent nothing; tool-driven turns (like questions) never trigger it.
  • Thrash guard: repeated identical tool failures stop the loop.

Word & Excel specifics ​

  • Selection context: text you highlight before typing is passed as <user_selection> and beats the whole-document state for ambiguous requests ("fix this", "make it bold").
  • Paragraph markers: <doc_state> marks headings, lists, page breaks and images so the model understands structure, not just text.

PowerPoint specifics ​

  • No selection API: PowerPoint.js has none, so there's never a <user_selection> block; ambiguous slide references are asked or resolved to the most recently used slide.
  • Re-import caveat: OOXML write-backs (edit_slide_text, edit_slide_xml) re-import the slide and reassign slide/shape ids; the tool returns a warning to re-run list_slide_shapes before further edits.
  • Stacking order (z-order): list_slide_shapes reports each shape's order (0 is the back, highest is the front). Full-bleed shapes added with insert_slide_element are automatically placed at the back, so backgrounds never cover text; pass z_order: "back" to force it for other shapes. verify_slides flags z_order_warnings when a background is stacked in front of content it hides.
  • Verification: verify_slides checks overlaps, out-of-bounds shapes, WCAG text contrast, and z-order problems after styling.

Team features on a self-hosted instance ​

Self-hosting adds two features aimed at teams:

  • SSO: users sign in with your organization's identity provider (OIDC, such as Microsoft Entra ID) before the add-in loads. See SSO.
  • Managed configuration: fix a subset of the settings, such as the provider and API key, for everyone on the instance. Everything else stays editable. See Managed configuration.

Both are included in the free self-hosted product. See Self-hosting to set it up.