Skip to content

AI chat and context

The AI machinery under apps/case/ai/ turns a matter's record into a system prompt, runs a model call on a background thread, reports progress to a polling view, and applies the writes the model was directed to make. Two surfaces share it: the case chat (the matter's AI tab) and the intake chat (apps/intakes/chat.py). Agentic-mode internals are on Agentic chat. All of it is optional: with no provider key configured, none of it is reachable (see AI is optional).

Where the code is

Module Holds
apps/case/ai/models.py Conversation, Message, MaterialChunk
apps/case/ai/views.py The case chat views: list, send, status poll, cancel, Compose Prompt, context preview
apps/case/ai/tasks.py process_ai_request() (the classic turn), finalize_response(), armed_write_protocols(), the model dispatch tables, the window-fit guard
apps/case/ai/context.py The context builders: section formatters, collect_context_items(), assemble_matter_context_with_selection(), build_request_info(), build_chat_history()
apps/case/ai/selector.py The context selector (fast tier): build_manifest(), select_context(), token estimates and model limits
apps/case/ai/prompts/legal.md The legal system prompt
apps/case/ai/status.py, access.py The cross-process status store and heartbeat; who may use a conversation and the Financial and Research gates
apps/case/ai/anthropic_client.py, gemini_client.py Provider calls, streaming, prompt caching, the tool loops
apps/case/ai/providers.py complete(), the one call for AI work that is not a matter chat; the fast and deep tiers and their model ids; provider(), chat_llm(), AINotConfigured
apps/settings/ai.py Whether AI is set up: configured_providers(), ai_enabled(), gemini_key(), anthropic_key(), key_source(), and the encryption of keys stored on Firm
config/context.py The integrations context processor: ai_enabled and caselaw_available for templates
apps/case/ai/fact_blocks.py, witness_blocks.py, note_blocks.py, caselaw_blocks.py The fenced write blocks
apps/case/ai/handles.py, citations.py, vetting.py Leaked [doc:ID]-style handles become links or note chips; citation extraction and CourtListener verification; the fast-tier vetting job
apps/case/ai/purge.py The closed-matter chat purge
apps/case/ai/semantic.py, embeddings.py, pricing.py The pgvector index behind the agent's search_materials; estimated cost for the status bar
templates/case/ai/ The chat window, status.html (the poller), prompt-editor-modal.html, message partials
apps/intakes/chat.py The intake chat, built on the same models and status protocol

Data model

A Conversation belongs to exactly one of a matter or an intake (both CASCADE); the code, not the database, keeps that exclusive. user is who started it (SET_NULL). llm and kind (classic, agent, or the retired research) are fixed when the conversation is created. ai_context (auto, always, never) says whether this conversation is offered to other conversations on the matter as a reference, and summary is the fast-tier précis the selector reads. vet_citations turns the post-answer vetting job on.

A Message has a role, content, the sending user for user messages, token counts, and JSON fields the renderers read: verified_citations, activity_log (the live log preserved on the answer), agent_run (agent turns) and research_trail (old research-kind answers). Both models keep HistoricalRecords.

MaterialChunk is one embedded chunk of a document, note, library note, email, highlight or fact, unique on (kind, object_id, chunk_index), with an HNSW index on its vector; semantic.py writes it on save through qcluster and manage.py build_semantic_index backfills.

How a turn runs

  1. Send. views.send_message() creates the conversation on the first message (title, llm, kind from the form), stores the user Message, seeds ai_status_<conversation id> with starting, and starts a daemon threading.Thread on tasks.process_ai_request(). The response is messages.html with the poller in it. The intake chat has its own send view and worker (_process_intake_chat()) on the same protocol.
  2. Context. For a classic matter chat the thread calls context.assemble_matter_context_with_selection() (below), appends a linked draft's section when the conversation has one, then SOURCE_LINKING (the citing conventions) and whatever write protocols armed_write_protocols() returns. An agent conversation branches here to agent.run_agent_request() instead.
  3. History and fit. build_chat_history() prefixes every message with its sent time and, when more than one person has taken part, the sender's name. fit_prompt_to_window() drops the oldest messages until the estimate fits 80% of the model window; for a Claude model whose estimate is past half the window it asks the API for an exact count and trims again against 98%. A context that cannot fit on its own raises PromptTooLargeError, which becomes the chat's error message.
  4. Model call. GEMINI_MODELS keys go to gemini_client.send_to_gemini_streaming() with thought summaries feeding the status line; everything else goes to anthropic_client.send_to_claude(). Fable keys get a higher output ceiling because their thinking cannot be turned off.
  5. Finalize. finalize_response() applies draft edits, the fenced write blocks, leaked-handle links, and the CourtListener citation check, in that order. It is the one path both modes write through.
  6. Complete. The thread writes a complete payload (response, tokens, citations, log) under the status key with FINAL_TTL, and the next poll stores the assistant Message, starts the summary and vetting threads, and deletes the key.

Compose Prompt (views.prompt_editor_modal(), templates/case/ai/prompt-editor-modal.html) is a rich-text editor whose markdown is posted to the same ai-send route; its hidden kind field carries the mode the window was opened in, because the first message creates the conversation. The draft lives in localStorage per conversation.

What goes into the context

assemble_matter_context_with_selection() builds the system prompt in this order: build_request_info() (today's date, the requesting user, the firm roster with titles, so the model never guesses a colleague's role), the legal prompt, then MATTER_CONTEXT_TEMPLATE: overview, contacts, witnesses, proceedings, the always-included items by importance tier, tasks, events, time entries and settlement. After the template come the selector's Selected Materials and an Also Available listing of what it left out, so the model can name an item and ask for it.

collect_context_items() gathers the always-included items: documents, saved cases and email threads with ai_context="always", every highlight and fact, and other conversations on the matter set to Always. Notes have had no AI setting since 2026-08-11 and are all selector material. Items set to Never are excluded everywhere.

format_time_entries(matter, include_billing=...) lists the work done; with include_billing each entry also carries its rate, fee, comp flag and invoice status. Since 2026-10-02 every requesting user gets that detail (the Activity screens show it to anyone who can see the matter). Invoices are different: build_manifest(include_invoices=...) offers them to the selector only when access.has_financial_access(user) is true. A run with no user sees neither money on time entries nor any invoice.

Library notes (standalone notes in folders flagged as AI library, apps.notes.models.get_library_notes()) are offered on every matter when include_library is true; the selector prompt tells the model to pick at most five and only when the topic bears on the question.

Reuse. The assembled context is cached under ai_ctx_<conversation id> in the cross-process ai_status cache (zlib-compressed, since it is the whole system prompt) for CONTEXT_REUSE_SECONDS (600) together with a fingerprint from _context_fingerprint(): count and latest updated_at per source table, the model, the user and their Financial flag. A follow-up inside ten minutes with an unchanged fingerprint skips the selector entirely, which also keeps the provider prompt caches warm. Any write to the material, including the AI's own note, fact or witness writes, changes the fingerprint. Tasks, events and time entries are deliberately not fingerprinted.

Ceilings. The selector works to MODEL_CONTEXT_LIMITS (about 60% of each window) less the fixed sections and a 10k reserve. If the whole prompt still estimates past 80% of MODEL_HARD_LIMITS, the selected items are demoted to the Also Available list; if the always-included content alone is still over, tiers are shed reference first, then medium, then high (critical items are never dropped) and listed under Omitted Materials. estimate_tokens() uses 2.5 characters per token: the folk figure of 4 let a prompt through that the provider rejected.

The selector

selector.build_manifest() lists every candidate as a ManifestItem (type, id, name, category, date, a description, word count, importance) with a content_map of the full text, resolved lazily for invoices and case law (the opinion is fetched from CourtListener on selection). The manifest covers auto documents with finished OCR, all matter notes, auto saved cases, other auto conversations (described by their summary), auto email threads as one item per thread, invoices when permitted, and library notes. With include_always=True (the agent's index) the always items are listed too, flagged pinned, each with the handle the agent tools read by.

select_context() skips the model call when the matter items total under SMALL_MATTER_THRESHOLD (30,000 words) and there are no invoices or library items: everything goes in. Otherwise SELECTOR_SYSTEM_PROMPT and the formatted manifest go to providers.complete() at the fast tier, which returns a JSON list of {type, id}; if the call or the parse fails, _fallback_by_importance() fills the budget by importance. Invoices and library notes never ride the short-circuit or the fallback: they enter only when the selector names them.

The system prompt file

apps/case/ai/prompts/legal.md is read fresh on every call by context.load_legal_prompt() (an edit takes effect without a restart) and its [JURISDICTION] placeholders are replaced with the matter's jurisdiction, else the firm's, else "United States common law". The same text heads the agent's orientation. The operator page AI providers and research describes what to edit in it.

Providers

The picker keys in Conversation.LLM_CHOICES map to provider model ids in tasks.CLAUDE_MODELS and tasks.GEMINI_MODELS; retired keys stay in the tables so old conversations still send, and views.available_llm_choices() offers only the models whose provider is in configured_providers(). The clients read their key through apps.settings.ai.gemini_key() and anthropic_key(), never from settings directly: the ANTHROPIC_API_KEY and GEMINI_API_KEY settings from config/.env (see the environment reference) first, then a key an admin stored under Settings > Integrations. Every model in both tables has a one-million-token window. Each client marks the system prompt cacheable (Anthropic's cache_control; a Gemini cachedContents object held ten minutes for prompts over 130k characters) and streams so a cancel stops the bill. Fable keys carry Anthropic's server-side refusal fallback to Opus 5 (anthropic_client.FALLBACK_MODELS).

Everything else goes through providers.complete(system, messages, tier): the selector, conversation, document, note and case summaries, citation vetting, intake extraction and assessment, the intake chat and AI quick-add. A tier maps to one model per provider (MODELS: fast is gemini-2.5-flash or claude-sonnet-4-6, deep is gemini-pro-latest or claude-sonnet-5), and provider(prefer) picks the preferred provider when it has a key, else the first configured one, Gemini first. So either key alone runs every feature. The intake chat stores chat_llm(DEEP) as its Conversation.llm and passes the matching prefer on each turn, so a chat stays on the provider it started on; quick-add passes the firm's model choice as prefer. With no provider, complete() raises AINotConfigured, which callers never reach in practice because the gates below stop them first.

Embeddings are the exception: embeddings.py is Gemini-only. Without gemini_key(), semantic._enqueue() queues nothing and semantic_entries() returns an empty list, so search is keyword-only with no error.

AI is optional

apps.settings.ai.ai_enabled() is true when configured_providers() is non-empty. A provider is configured when its key is in config/.env or stored on the Firm row (gemini_api_key, anthropic_api_key). Stored keys are Fernet tokens under a key derived from SECRET_KEY (_fernet()), so a new SECRET_KEY makes decrypt_key() return "" with a logged warning, and the provider counts as unconfigured. The admin-only form is ai_key_save in apps/settings/integrations/views.py (templates/settings/integrations/ai.html): verify_ai_key() lists the provider's models with the candidate key before saving it, and an empty value removes it. There is no cache: each check reads the Firm row (at most once per call), so a saved key takes effect in every process at once.

Four layers keep a server without a key free of AI:

  • Templates. The config.context.integrations context processor puts lazy ai_enabled and caselaw_available in every template context (evaluated only when a template reads them). They hide the AI tab and its Case Law switch (case-nav.html, case/ai/view-pills.html), the intake Assessment pill and its Assess and Chat buttons, the AI column and bulk menu on Documents and Case Law, and the Settings > Tasks menu entry.
  • Routes. PermissionMiddleware.__call__ matches AI_PATTERN (/case/…/ai/, /case/drafts/, /intakes/<id>/assess, /intakes/<id>/chat/ and /settings/tasks/) and answers 404 for everyone, admins included, when ai_enabled() is false. CASELAW_PATTERN does the same for saved case law and the cluster viewer when there is no CourtListener token. Both run before the permission checks.
  • The remembered tab. tab_available(user, tab) in apps/case/views.py says whether ai or caselaws can be shown; get_last_tab() falls back to Documents for a stored tab that no longer is.
  • Background work. Every summary task, its queueing helper, the semantic enqueue and the inbound intake pipeline check ai_enabled() (or gemini_key()) first and return. FilesForm drops ai_context, leaving the stored value alone. quick_add_ai_enabled() requires both Firm.quick_task_ai and ai_enabled().

Saved case law follows CourtListener rather than AI: courtlistener.caselaw_available(user) is a token plus admin or perm_research. With AI on, the list is a view of the AI tab; with AI off, case-nav.html shows it as its own Case Law tab.

Tests: the root conftest.py gives every test fake Gemini, Anthropic and CourtListener keys (autouse _integration_keys), the same on a laptop with real keys in config/.env as in CI, so a stray real call fails instead of spending. Request the ai_off or courtlistener_off fixture to test the unset case; apps/case/tests/test_ai_optional.py and apps/settings/tests/test_ai_keys.py are the examples.

Status and the poller

status.py holds the run status under ai_status_<conversation id> in the ai_status cache, a DatabaseCache on the ai_status_cache table (config/settings.py; manage.py createcachetable creates it, not a migration). It replaced the per-process LocMemCache on 2026-08-17: production runs several gunicorn workers, a poll usually lands in a different worker from the run thread, and each worker's private view made most first polls find nothing and fabricate "server restarted" replies.

Liveness is TTL-based. In-flight writes carry RUNNING_TTL (180 s) and a RunHeartbeat thread re-touches the entry every HEARTBEAT_SECONDS (30). Terminal payloads carry FINAL_TTL (600 s) so a poller that arrives late can still collect them. views.ai_status() is shared by the two chats: a missing key while the latest message is still the user's means the process died, and the view writes the "interrupted" assistant message itself; complete and error are stored as messages; cancelled is left in place so the thread's is_cancelled() keeps seeing it.

templates/case/ai/status.html polls every second with hx-swap="morph" (idiomorph) so the box is patched, not replaced; a page that hosts it must load idiomorph. The interval ends when a terminal reply swaps in a different root element, and views._terminal() adds HX-Reswap: outerHTML so that swap also works without idiomorph: before it, the intake window's poller ran on forever and wrote an "interrupted" message every second (2026-08-23).

Fenced write blocks

The model writes to the record only through fenced blocks in its reply: create-facts and create-witnesses (lists), create-note and edit-note (one object each) and save-caselaw (agent conversations only). The intake chat has its own update-intake block. A protocol is added to the system prompt only when the last few user messages (the current one and the three before it) match its trigger pattern (FACTS_TRIGGER_RE and the like), so an ordinary chat carries no standing write instructions to misfire on, and a follow-up ("also add the crash date") keeps the protocol armed.

finalize_response() applies each block with a regex substitution over the reply (apply_fact_blocks(), apply_witness_blocks(), apply_caselaw_blocks(), apply_note_blocks()), replacing the block with confirmation lines so the user and the model's later turns both see what happened. Nothing is confirmed first: the row exists by the time the answer renders. A block that is not valid JSON is left as text and writes nothing. Entry creators are forgiving on optional fields and strict on the row itself (a fact's description, a witness's name); cited document and highlight ids are filtered to the matter, so a guessed id links nothing; a same-named witness is reused; a case already saved gets the proposition appended to its notes rather than a duplicate row ((matter, cluster_id) is unique). Each creator is also what the MCP server's write tools call (see MCP).

Note writes apply on the chat's worker thread, synchronously, and are never refused: the per-note Note.ai_write_until grant gates only the MCP server's write_note, not the in-app AI. Creation is limited to the matter's notes; edits reach the matter's notes and library notes, and every edit lands in the note's history under the requesting user, so a bad one is recoverable. Because a model that has seen real confirmations sometimes imitates the line instead of emitting the block, strip_fake_note_confirmations() runs on the raw reply before the blocks are applied and turns a lookalike into an honest "no note was changed" notice.

Background work

Chat turns are not qcluster jobs: they run on daemon threads inside the web worker, so a deploy or a .py reload kills them, the heartbeat stops, the key expires and the next poll reports the interruption. So do the conversation summary (tasks.generate_conversation_summary()) and the vetting job, started from the status view.

purge.purge_closed_chats() deletes every matter conversation, its messages and their history rows once the matter's current Closed streak (read from the matter's history) is older than CHAT_RETENTION_DAYS (180; 0 keeps chats), from the weekly chat-purge-weekly schedule or the purge_closed_chats command. Intake chats are not in scope: they are deleted when ended or discarded.

Access

Routes under /case/<matter id>/... are checked for matter membership by PermissionMiddleware.process_view() in apps/accounts/middleware.py. That check never sees a conversation id in the query string or the body, and the status, cancel and intake routes name no matter, so access.py fills the gaps: conversation_for_user() (a matter chat needs the matter, an intake chat the Intakes permission) guards the poll and cancel views, and matter_conversation_for_user() ties a posted conversation id to the matter in the URL with a 404, so a wrong id confirms nothing.

The reply is built for the user who asks (tightened 2026-10-02): has_financial_access() (admin or perm_financial) decides whether invoices are offered, and has_research_access() (admin or perm_research) whether the agent gets the case-law tools and the save protocol. The matrix is in the permissions reference.

Things that bite

  • Never edit a .py while a run may be in flight. The worker reloads, the daemon thread dies with it, and the user gets the "interrupted" reply. Agent runs are minutes long.
  • Order inside finalize_response() is fixed. Draft edits first; strip_fake_note_confirmations() before apply_note_blocks() (after, a real confirmation would match); handles after the blocks that consume them; citations last.
  • The model tables must stay in step. A new picker key needs LLM_CHOICES (and a migration for the choices), CLAUDE_MODELS or GEMINI_MODELS, MODEL_CONTEXT_LIMITS, MODEL_HARD_LIMITS and the pricing.py rates. Anthropic publishes no "latest" alias, so each version is a new key.
  • The intake chat tests monkeypatch threading.Thread to run the worker inline; status.py binds Thread at import so the heartbeat stays a real thread, or that patch turns its wait loop into a hang.