Backup and restore¶
For the person running a Kosmos server. This page lists everything that holds a firm's data, and gives a procedure for backing it up and for restoring it.
Kosmos does not back itself up
The repository contains no backup tooling: no script, no scheduled job, no management command. Nothing is backed up unless you set it up. The procedures on this page are standard PostgreSQL and file-copy practice applied to what the code stores. They are not run by the project's automated tests, so rehearse a restore yourself before you rely on them.
What holds state¶
| What | Where | Notes |
|---|---|---|
| The database | PostgreSQL, the database named by DB_NAME in config/.env (the installer's default is kosmos) |
Almost everything: matters, contacts, time and billing, trust ledgers, notes, AI conversations, emails synced from Gmail, extracted document text, the search indexes, the task queue and schedules, user accounts and sessions. |
| Uploaded files, local storage | media/ in the checkout, when STORAGE_BACKEND=local |
Documents (media/documents/), stored invoice PDFs (media/invoices/) and the firm logo (media/company/). |
| Uploaded files, object storage | The bucket named by DIGITAL_OCEAN_BUCKET_NAME, when STORAGE_BACKEND=s3 |
The same files, under the same paths, in an S3-compatible bucket instead of media/. |
| Configuration | config/.env |
Holds SECRET_KEY, the database password and every API key set there (AI keys entered under Settings → Integrations are in the database, encrypted with SECRET_KEY). It is not in git. |
| Google credentials | The directory named by GOOGLE_DATA_DIR (default google/ in the checkout) |
The OAuth client file you downloaded from Google (google_tokens.json) and the tokens Kosmos obtained for Calendar, Contacts and Drive. Not in git. |
| The code version | The git commit of the checkout | Record it with each backup: git rev-parse HEAD. |
The database and the uploaded files belong together. A document is a row in the database plus a file in storage, and neither is useful without the other.
Files that are yours but can be rebuilt, worth a copy if you have changed them:
gunicorn.conf.pyin the checkout, if you edited it.- The installed systemd units and nginx files, if you edited them, and
/etc/letsencrypt/if you would rather not have certbot issue a new certificate after a rebuild.
Things you do not need to back up:
.venv/. The installer recreates it.logs/. Keep them if your firm's policy calls for it. Nothing reads them back..dev-sessions/. It exists only whenENV=dev, where login sessions are kept in files. On a production install sessions are in the database.
Why config/.env matters¶
Keep the same SECRET_KEY across a restore. Payment links and client
intake-form links in emails the firm has already sent are signed with it,
so a restored server with a new key rejects those links. A new key also
signs every user out, and makes AI keys entered under Settings →
Integrations unreadable, so they have to be entered again.
Treat backups as confidential¶
A backup contains privileged client material and live credentials: the
API keys in config/.env, the Google tokens, and the Gmail access tokens
of users who connected a mailbox, which are stored in the database.
Encrypt backups, restrict who can read them, and store a copy off the
server.
Back up¶
Do these in order, database first. A file added between the two steps is then merely absent from the database copy, which is harmless. The other order can leave database rows pointing at files the backup does not have.
-
Record the code version.
-
Dump the database. Use the name, role and password from
config/.env:pg_dumptakes a consistent snapshot while the application keeps running, so there is no need to stop anything. It asks for the role's password. For an unattended job, put the password in the account's~/.pgpassfile, as described in the PostgreSQL documentation. -
Copy the uploaded files.
With
STORAGE_BACKEND=local:With
STORAGE_BACKEND=s3, copy the bucket with your storage provider's tools, or any S3 client. For example, with the AWS command line client and the values fromconfig/.env: -
Copy the configuration and the Google credentials.
-
Move the result off the server, encrypted.
Two things to decide for yourself, because the project does not:
- How often, and how long to keep each backup. Run the steps above from cron or a systemd timer.
- How to keep deleted files recoverable. When a document is deleted
in Kosmos, its file is removed from storage at once. A backup that
mirrors storage exactly (
rsync --delete, or a bucket sync with deletion) therefore loses the file too. Keep dated copies, or switch on object versioning in the bucket, if you need to recover a file after a mistaken deletion.
Restore on the same server¶
Use this to return a working server to an earlier backup.
-
Stop the application and the worker, so nothing writes while you restore:
-
Replace the database with an empty one owned by the application's role, and create the two extensions the schema needs. This must be done as the
postgressuperuser:sudo -u postgres dropdb kosmos sudo -u postgres createdb -O kosmos kosmos sudo -u postgres psql -d kosmos \ -c 'CREATE EXTENSION IF NOT EXISTS pg_trgm' \ -c 'CREATE EXTENSION IF NOT EXISTS vector'If the second extension fails with
could not open extension control file, the pgvector package for this PostgreSQL version is not installed. -
Load the dump:
This expects the role that owned the tables when the dump was taken to exist under the same name. If your role is named differently now, add
--no-owner --role=YOUR_ROLE. -
Put the files back. With local storage:
With object storage, copy the files back into the bucket.
-
Put
config/.envand the Google credential directory back, if they were lost or changed. -
Run the post-restore commands, as the application's account:
.venv/bin/python manage.py migrate --noinput .venv/bin/python manage.py createcachetable .venv/bin/python manage.py installwatson .venv/bin/python manage.py buildwatson .venv/bin/python manage.py setup_schedulesCommand Why migrateBrings the schema up to the code in the checkout, if the checkout is newer than the backup. It does nothing otherwise. createcachetableMakes sure the ai_status_cachetable exists. AI chat needs it and no migration creates it.installwatson,buildwatsonMake sure the full-text search column exists, and rebuild the search index. setup_schedulesBrings the scheduled jobs up to date and moves each one to its next slot. Without this, every job that looks overdue in the restored data runs as soon as the worker starts. The checkout must not be older than the backup. Old code cannot run against a newer schema.
-
Start everything:
Restore onto a new server¶
-
Clone the repository and check out the commit recorded with the backup, or a newer one.
-
Put the backed-up
config/.envin place before installing, with mode600:The installer keeps an existing
config/.envand takes the database name, role and password from it. -
Run the installer. See Install with the installer script.
This installs PostgreSQL, pgvector and the other packages, and creates the role and an empty database.
-
Follow Restore on the same server from step 1 to the end. Step 2 there replaces the empty database the installer just created.
-
If the hostname changed, update
ALLOWED_HOSTS,CSRF_TRUSTED_ORIGINSandPUBLIC_BASE_URLinconfig/.env, restart both services, and run certbot for the new name.
Check that it worked¶
systemctl is-active law.socket law.service qcluster.service
curl -fsS https://kosmos.example.com/health/ready/
Then sign in and check the things a database-only test would miss:
- Open a document in a matter. This proves that the files and the database match.
- Search for a word you know appears in a document. This proves the search index was rebuilt.
-
List document records whose file is missing from storage:
The command only lists. It deletes nothing unless you pass
--apply, so do not pass it here. For files that were mirrored from Google Drive,restore_drive_documentscan download missing ones again (it only reports, unless you pass--apply).
Rehearsing a restore¶
Restore to a spare machine from time to time. A backup that has never been restored is a guess.
A rehearsal copy is a complete, live copy of the firm's system. Before
starting it, edit its config/.env so that it cannot act on the outside
world:
- set
ENVtodev, so that every page is marked as a development instance; - set
EMAIL_BACKEND=console, so that nothing is emailed to users or clients (login codes are then printed tologs/error.log, as on a fresh install); - set
PAYMENT_PROCESSOR=fakeand blank the processor keys; - do not copy the Google credential directory, so the firm-wide Calendar, Contacts and Drive connections stay off.
The background worker is what sends the daily digest and runs the sync
jobs. If the rehearsal only needs to prove that the data is there, leave
qcluster.service stopped. See Background worker.