AI chat and context¶
The AI machinery under apps/case/ai/ turns a matter's record into a
system prompt, runs a model call on a background thread, reports progress
to a polling view, and applies the writes the model was directed to make.
Two surfaces share it: the case chat (the matter's AI tab) and the
intake chat (apps/intakes/chat.py). Agentic-mode internals are on
Agentic chat. All of it is optional: with no provider
key configured, none of it is reachable (see
AI is optional).
Where the code is¶
| Module | Holds |
|---|---|
apps/case/ai/models.py |
Conversation, Message, MaterialChunk |
apps/case/ai/views.py |
The case chat views: list, send, status poll, cancel, Compose Prompt, context preview |
apps/case/ai/tasks.py |
process_ai_request() (the classic turn), finalize_response(), armed_write_protocols(), the model dispatch tables, the window-fit guard |
apps/case/ai/context.py |
The context builders: section formatters, collect_context_items(), assemble_matter_context_with_selection(), build_request_info(), build_chat_history() |
apps/case/ai/selector.py |
The context selector (fast tier): build_manifest(), select_context(), token estimates and model limits |
apps/case/ai/prompts/legal.md |
The legal system prompt |
apps/case/ai/status.py, access.py |
The cross-process status store and heartbeat; who may use a conversation and the Financial and Research gates |
apps/case/ai/anthropic_client.py, gemini_client.py |
Provider calls, streaming, prompt caching, the tool loops |
apps/case/ai/providers.py |
complete(), the one call for AI work that is not a matter chat; the fast and deep tiers and their model ids; provider(), chat_llm(), AINotConfigured |
apps/settings/ai.py |
Whether AI is set up: configured_providers(), ai_enabled(), gemini_key(), anthropic_key(), key_source(), and the encryption of keys stored on Firm |
config/context.py |
The integrations context processor: ai_enabled and caselaw_available for templates |
apps/case/ai/fact_blocks.py, witness_blocks.py, note_blocks.py, caselaw_blocks.py |
The fenced write blocks |
apps/case/ai/handles.py, citations.py, vetting.py |
Leaked [doc:ID]-style handles become links or note chips; citation extraction and CourtListener verification; the fast-tier vetting job |
apps/case/ai/purge.py |
The closed-matter chat purge |
apps/case/ai/semantic.py, embeddings.py, pricing.py |
The pgvector index behind the agent's search_materials; estimated cost for the status bar |
templates/case/ai/ |
The chat window, status.html (the poller), prompt-editor-modal.html, message partials |
apps/intakes/chat.py |
The intake chat, built on the same models and status protocol |
Data model¶
A Conversation belongs to exactly one of a matter or an intake
(both CASCADE); the code, not the database, keeps that exclusive.
user is who started it (SET_NULL). llm and kind (classic, agent, or
the retired research) are fixed when the conversation is created.
ai_context (auto, always, never) says whether this conversation is
offered to other conversations on the matter as a reference, and summary
is the fast-tier précis the selector reads. vet_citations turns the
post-answer vetting job on.
A Message has a role, content, the sending user for user messages,
token counts, and JSON fields the renderers read: verified_citations,
activity_log (the live log preserved on the answer), agent_run (agent
turns) and research_trail (old research-kind answers). Both models keep
HistoricalRecords.
MaterialChunk is one embedded chunk of a document, note, library note,
email, highlight or fact, unique on (kind, object_id, chunk_index), with
an HNSW index on its vector; semantic.py writes it on save through
qcluster and manage.py build_semantic_index backfills.
How a turn runs¶
- Send.
views.send_message()creates the conversation on the first message (title,llm,kindfrom the form), stores the userMessage, seedsai_status_<conversation id>withstarting, and starts a daemonthreading.Threadontasks.process_ai_request(). The response ismessages.htmlwith the poller in it. The intake chat has its own send view and worker (_process_intake_chat()) on the same protocol. - Context. For a classic matter chat the thread calls
context.assemble_matter_context_with_selection()(below), appends a linked draft's section when the conversation has one, thenSOURCE_LINKING(the citing conventions) and whatever write protocolsarmed_write_protocols()returns. An agent conversation branches here toagent.run_agent_request()instead. - History and fit.
build_chat_history()prefixes every message with its sent time and, when more than one person has taken part, the sender's name.fit_prompt_to_window()drops the oldest messages until the estimate fits 80% of the model window; for a Claude model whose estimate is past half the window it asks the API for an exact count and trims again against 98%. A context that cannot fit on its own raisesPromptTooLargeError, which becomes the chat's error message. - Model call.
GEMINI_MODELSkeys go togemini_client.send_to_gemini_streaming()with thought summaries feeding the status line; everything else goes toanthropic_client.send_to_claude(). Fable keys get a higher output ceiling because their thinking cannot be turned off. - Finalize.
finalize_response()applies draft edits, the fenced write blocks, leaked-handle links, and the CourtListener citation check, in that order. It is the one path both modes write through. - Complete. The thread writes a
completepayload (response, tokens, citations, log) under the status key withFINAL_TTL, and the next poll stores the assistantMessage, starts the summary and vetting threads, and deletes the key.
Compose Prompt (views.prompt_editor_modal(),
templates/case/ai/prompt-editor-modal.html) is a rich-text editor whose
markdown is posted to the same ai-send route; its hidden kind field
carries the mode the window was opened in, because the first message
creates the conversation. The draft lives in localStorage per conversation.
What goes into the context¶
assemble_matter_context_with_selection() builds the system prompt in
this order: build_request_info() (today's date, the requesting user,
the firm roster with titles, so the model never guesses a colleague's
role), the legal prompt, then MATTER_CONTEXT_TEMPLATE: overview,
contacts, witnesses, proceedings, the always-included items by importance
tier, tasks, events, time entries and settlement. After the template come
the selector's Selected Materials and an Also Available listing of
what it left out, so the model can name an item and ask for it.
collect_context_items() gathers the always-included items: documents,
saved cases and email threads with ai_context="always", every highlight
and fact, and other conversations on the matter set to Always. Notes
have had no AI setting since 2026-08-11 and are all selector material.
Items set to Never are excluded everywhere.
format_time_entries(matter, include_billing=...) lists the work done;
with include_billing each entry also carries its rate, fee, comp flag
and invoice status. Since 2026-10-02 every requesting user gets that
detail (the Activity screens show it to anyone who can see the matter).
Invoices are different: build_manifest(include_invoices=...) offers them
to the selector only when access.has_financial_access(user) is true. A
run with no user sees neither money on time entries nor any invoice.
Library notes (standalone notes in folders flagged as AI library,
apps.notes.models.get_library_notes()) are offered on every matter when
include_library is true; the selector prompt tells the model to pick at
most five and only when the topic bears on the question.
Reuse. The assembled context is cached under ai_ctx_<conversation id>
in the cross-process ai_status cache (zlib-compressed, since it is the
whole system prompt) for CONTEXT_REUSE_SECONDS (600) together with a
fingerprint from _context_fingerprint(): count and latest updated_at
per source table, the model, the user and their Financial flag. A follow-up inside ten
minutes with an unchanged fingerprint skips the selector entirely, which
also keeps the provider prompt caches warm. Any write to the material,
including the AI's own note, fact or witness writes, changes the
fingerprint. Tasks, events and time entries are deliberately not
fingerprinted.
Ceilings. The selector works to MODEL_CONTEXT_LIMITS (about 60% of
each window) less the fixed sections and a 10k reserve. If the whole
prompt still estimates past 80% of MODEL_HARD_LIMITS, the selected
items are demoted to the Also Available list; if the always-included
content alone is still over, tiers are shed reference first, then
medium, then high (critical items are never dropped) and listed under
Omitted Materials. estimate_tokens() uses 2.5 characters per token:
the folk figure of 4 let a prompt through that the provider rejected.
The selector¶
selector.build_manifest() lists every candidate as a ManifestItem
(type, id, name, category, date, a description, word count, importance)
with a content_map of the full text, resolved lazily for invoices and
case law (the opinion is fetched from CourtListener on selection). The
manifest covers auto documents with finished OCR, all matter notes,
auto saved cases, other auto conversations (described by their
summary), auto email threads as one item per thread, invoices when
permitted, and library notes. With include_always=True (the agent's
index) the always items are listed too, flagged pinned, each with the
handle the agent tools read by.
select_context() skips the model call when the matter items total under
SMALL_MATTER_THRESHOLD (30,000 words) and there are no invoices or
library items: everything goes in. Otherwise SELECTOR_SYSTEM_PROMPT and
the formatted manifest go to providers.complete() at the fast tier,
which returns a JSON list
of {type, id}; if the call or the parse fails, _fallback_by_importance()
fills the budget by importance. Invoices and library notes never ride the
short-circuit or the fallback: they enter only when the selector names them.
The system prompt file¶
apps/case/ai/prompts/legal.md is read fresh on every call by
context.load_legal_prompt() (an edit takes effect without a restart)
and its [JURISDICTION] placeholders are replaced with the matter's
jurisdiction, else the firm's, else "United States common law". The same
text heads the agent's orientation. The operator page
AI providers and research describes
what to edit in it.
Providers¶
The picker keys in Conversation.LLM_CHOICES map to provider model ids
in tasks.CLAUDE_MODELS and tasks.GEMINI_MODELS; retired keys stay in
the tables so old conversations still send, and
views.available_llm_choices() offers only the models whose provider is
in configured_providers(). The clients read their key through
apps.settings.ai.gemini_key() and anthropic_key(), never from
settings directly: the ANTHROPIC_API_KEY and GEMINI_API_KEY
settings from config/.env (see the
environment reference) first, then
a key an admin stored under Settings > Integrations. Every model in
both tables has a one-million-token window. Each client marks the system
prompt cacheable (Anthropic's cache_control; a Gemini cachedContents
object held ten minutes for prompts over 130k characters) and streams so
a cancel stops the bill. Fable keys carry Anthropic's server-side refusal
fallback to Opus 5 (anthropic_client.FALLBACK_MODELS).
Everything else goes through providers.complete(system, messages,
tier): the selector, conversation, document, note and case summaries,
citation vetting, intake extraction and assessment, the intake chat and
AI quick-add. A tier maps to one model per provider (MODELS: fast is
gemini-2.5-flash or claude-sonnet-4-6, deep is gemini-pro-latest
or claude-sonnet-5), and provider(prefer) picks the preferred
provider when it has a key, else the first configured one, Gemini first.
So either key alone runs every feature. The intake chat stores
chat_llm(DEEP) as its Conversation.llm and passes the matching
prefer on each turn, so a chat stays on the provider it started on;
quick-add passes the firm's model choice as prefer. With no provider,
complete() raises AINotConfigured, which callers never reach in
practice because the gates below stop them first.
Embeddings are the exception: embeddings.py is Gemini-only. Without
gemini_key(), semantic._enqueue() queues nothing and
semantic_entries() returns an empty list, so search is keyword-only
with no error.
AI is optional¶
apps.settings.ai.ai_enabled() is true when configured_providers() is
non-empty. A provider is configured when its key is in config/.env or
stored on the Firm row (gemini_api_key, anthropic_api_key). Stored
keys are Fernet tokens under a key derived from SECRET_KEY
(_fernet()), so a new SECRET_KEY makes decrypt_key() return ""
with a logged warning, and the provider counts as unconfigured. The
admin-only form is ai_key_save in apps/settings/integrations/views.py
(templates/settings/integrations/ai.html): verify_ai_key() lists the
provider's models with the candidate key before saving it, and an empty
value removes it. There is no cache: each check reads the
Firm row (at most once per call), so a saved key takes effect in every
process at once.
Four layers keep a server without a key free of AI:
- Templates. The
config.context.integrationscontext processor puts lazyai_enabledandcaselaw_availablein every template context (evaluated only when a template reads them). They hide the AI tab and its Case Law switch (case-nav.html,case/ai/view-pills.html), the intake Assessment pill and its Assess and Chat buttons, the AI column and bulk menu on Documents and Case Law, and the Settings > Tasks menu entry. - Routes.
PermissionMiddleware.__call__matchesAI_PATTERN(/case/…/ai/,/case/drafts/,/intakes/<id>/assess,/intakes/<id>/chat/and/settings/tasks/) and answers 404 for everyone, admins included, whenai_enabled()is false.CASELAW_PATTERNdoes the same for saved case law and the cluster viewer when there is no CourtListener token. Both run before the permission checks. - The remembered tab.
tab_available(user, tab)inapps/case/views.pysays whetheraiorcaselawscan be shown;get_last_tab()falls back to Documents for a stored tab that no longer is. - Background work. Every summary task, its queueing helper, the
semantic enqueue and the inbound intake pipeline check
ai_enabled()(orgemini_key()) first and return.FilesFormdropsai_context, leaving the stored value alone.quick_add_ai_enabled()requires bothFirm.quick_task_aiandai_enabled().
Saved case law follows CourtListener rather than AI:
courtlistener.caselaw_available(user) is a token plus admin or
perm_research. With AI on, the list is a view of the AI tab; with AI
off, case-nav.html shows it as its own Case Law tab.
Tests: the root conftest.py gives every test fake Gemini, Anthropic and
CourtListener keys (autouse _integration_keys), the same on a laptop
with real keys in config/.env as in CI, so a stray real call fails
instead of spending. Request the ai_off or courtlistener_off fixture
to test the unset case; apps/case/tests/test_ai_optional.py and
apps/settings/tests/test_ai_keys.py are the examples.
Status and the poller¶
status.py holds the run status under ai_status_<conversation id> in
the ai_status cache, a DatabaseCache on the ai_status_cache table
(config/settings.py; manage.py createcachetable creates it, not a
migration). It replaced the per-process LocMemCache on 2026-08-17:
production runs several gunicorn workers, a poll usually lands in a
different worker from the run thread, and each worker's private view made
most first polls find nothing and fabricate "server restarted" replies.
Liveness is TTL-based. In-flight writes carry RUNNING_TTL (180 s) and a
RunHeartbeat thread re-touches the entry every HEARTBEAT_SECONDS
(30). Terminal payloads carry FINAL_TTL (600 s) so a poller that arrives
late can still collect them. views.ai_status() is shared by the two
chats: a missing key while the latest message is still the user's means
the process died, and the view writes the "interrupted" assistant message
itself; complete and error are stored as messages; cancelled is left
in place so the thread's is_cancelled() keeps seeing it.
templates/case/ai/status.html polls every second with hx-swap="morph"
(idiomorph) so the box is patched, not replaced; a page that hosts it
must load idiomorph. The interval ends when a terminal reply swaps in a
different root element, and views._terminal() adds HX-Reswap: outerHTML
so that swap also works without idiomorph: before it, the intake window's
poller ran on forever and wrote an "interrupted" message every second
(2026-08-23).
Fenced write blocks¶
The model writes to the record only through fenced blocks in its reply:
create-facts and create-witnesses (lists), create-note and
edit-note (one object each) and save-caselaw (agent conversations
only). The intake chat has its own update-intake block. A protocol is added to the system prompt
only when the last few user messages (the current one and the three
before it) match its trigger pattern (FACTS_TRIGGER_RE and the like),
so an ordinary chat carries no standing write instructions to misfire
on, and a follow-up ("also add the crash date") keeps the protocol armed.
finalize_response() applies each block with a regex substitution over
the reply (apply_fact_blocks(), apply_witness_blocks(),
apply_caselaw_blocks(), apply_note_blocks()), replacing the block with
confirmation lines so the user and the model's later turns both see what
happened. Nothing is confirmed first: the row exists by the time the
answer renders. A block that is not valid JSON is left as text and writes
nothing. Entry creators are forgiving on optional fields and strict on
the row itself (a fact's description, a witness's name); cited document
and highlight ids are filtered to the matter, so a guessed id links
nothing; a same-named witness is reused; a case already saved gets the
proposition appended to its notes rather than a duplicate row
((matter, cluster_id) is unique). Each creator is also what the MCP
server's write tools call (see MCP).
Note writes apply on the chat's worker thread, synchronously, and are
never refused: the per-note Note.ai_write_until grant gates only the
MCP server's write_note, not the in-app AI. Creation is limited to the
matter's notes; edits reach the matter's notes and library notes, and
every edit lands in the note's history under the requesting user, so a
bad one is recoverable. Because a model that has seen real confirmations
sometimes imitates the line instead of emitting the block,
strip_fake_note_confirmations() runs on the raw reply before the blocks
are applied and turns a lookalike into an honest "no note was changed"
notice.
Background work¶
Chat turns are not qcluster jobs: they run on daemon threads inside the
web worker, so a deploy or a .py reload kills them, the heartbeat
stops, the key expires and the next poll reports the interruption. So do
the conversation summary (tasks.generate_conversation_summary()) and
the vetting job, started from the status view.
purge.purge_closed_chats() deletes every matter conversation, its
messages and their history rows once the matter's current Closed streak
(read from the matter's history) is older than CHAT_RETENTION_DAYS
(180; 0 keeps chats), from the weekly chat-purge-weekly schedule or the
purge_closed_chats command. Intake chats are not in scope: they are
deleted when ended or discarded.
Access¶
Routes under /case/<matter id>/... are checked for matter membership by
PermissionMiddleware.process_view() in apps/accounts/middleware.py.
That check never sees a conversation id in the query string or the body,
and the status, cancel and intake routes name no matter, so access.py
fills the gaps: conversation_for_user() (a matter chat needs the
matter, an intake chat the Intakes permission) guards the poll and cancel views, and
matter_conversation_for_user() ties a posted conversation id to the
matter in the URL with a 404, so a wrong id confirms nothing.
The reply is built for the user who asks (tightened 2026-10-02):
has_financial_access() (admin or perm_financial) decides whether
invoices are offered, and has_research_access() (admin or
perm_research) whether the agent gets the case-law tools and the save
protocol. The matrix is in the
permissions reference.
Things that bite¶
- Never edit a
.pywhile a run may be in flight. The worker reloads, the daemon thread dies with it, and the user gets the "interrupted" reply. Agent runs are minutes long. - Order inside
finalize_response()is fixed. Draft edits first;strip_fake_note_confirmations()beforeapply_note_blocks()(after, a real confirmation would match); handles after the blocks that consume them; citations last. - The model tables must stay in step. A new picker key needs
LLM_CHOICES(and a migration for the choices),CLAUDE_MODELSorGEMINI_MODELS,MODEL_CONTEXT_LIMITS,MODEL_HARD_LIMITSand thepricing.pyrates. Anthropic publishes no "latest" alias, so each version is a new key. - The intake chat tests monkeypatch
threading.Threadto run the worker inline;status.pybindsThreadat import so the heartbeat stays a real thread, or that patch turns its wait loop into a hang.
Related¶
- AI chat in the user guide shows the screens.
- AI providers and research covers keys, the prompt file and retention for operators; Background worker covers qcluster.
- Neighbouring pages: Agentic chat, MCP server and JSON APIs, Case building, Notes and drafts, Email and intakes.