Skip to content

feat: Add treplica helper for PostgreSQL streaming read replicas - #404

Open
O1eq wants to merge 3 commits into
totara:masterfrom
O1eq:feat-add-treplica
Open

O1eq wants to merge 3 commits into
totara:masterfrom
O1eq:feat-add-treplica

Conversation

@O1eq

@O1eq O1eq commented Aug 7, 2026

Copy link
Copy Markdown

Description

Adds a new bin/treplica command that creates and manages a PostgreSQL streaming read replica of any totara pgsql container. Its purpose is testing Totara's database read/write splitting feature ($CFG->dboptions['readonly']), which routes read queries to a read-only replica.

  • treplica up [pgsqlNN] - clones the running primary with pg_basebackup into a new volume and starts it as a hot standby (default service: pgsql16). Image, PGDATA and resource settings (max_connections, max_locks_per_transaction, ...) are derived from the running primary, so it works with any pgsql version/config without hardcoded values. Idempotent: re-running shows status, a stopped replica is restarted, a stale volume is refused with a hint. Ends by printing the ready-to-paste config.php block (including a warning that it must be placed after the config-after.phpinclude, which reassigns $CFG->dboptions).
  • treplica status - recovery check, pg_stat_replication, replication lag.
  • treplica logs - follows the standby's statement log.
  • treplica down - removes container + volume after confirmation and drop the replication slot if one was created.
  • Options: --port (publish on a host port), --slot (physical replication slot for bulk-load safety), --quiet-logs.

Testing Instructions

  1. Check out this pull request
  2. Have any Totara site running on a pgsql container (e.g. pgsql16)
  3. Run treplica up - it should finish with state: streaming / Replica lag: in sync and print a config.php snippet
  4. Run treplica status - same healthy output; run treplica up again - it should no-op politely
  5. (Optional, needs a t21/TL-36188 site) paste the printed readonly block at the end of your config.php, set $CFG->perfdebug = 15;, reload any page and check the footer shows "DB reads from replica" counting up, while treplica logs shows only SELECT statements
  6. Run treplica down, confirm - container, volume (and slot, if --slot was used) are removed; treplica status now reports there is no replica

Checklist

  • Does what the author says it will do
  • Testing instructions are provided
  • Commit messages make sense and follow the conventional commit standard
  • No identified security issues
  • No identified maintenance issues
  • Any third-party libraries/dependencies use the MIT or Apache 2.0 license
  • Changes made are backwards compatible and will not break existing setups
  • Changes to scripts in the bin/ directory run correctly on both MacOS and WSL
  • Changes to containers can be built locally sucessfully (e.g. via tbuild container && tup container)
  • Containers/images are compatible with both AMD64 (Windows) and ARM64 (MacOS)
  • Changes made to config.php are compatible with our oldest supported Totara version, our newest Totara version, and Moodle

Comment thread bin/treplica
@O1eq O1eq changed the title Add treplica helper for PostgreSQL streaming read replicas feat: add treplica helper for PostgreSQL streaming read replicas Aug 17, 2026
@O1eq O1eq changed the title feat: add treplica helper for PostgreSQL streaming read replicas Add treplica helper for PostgreSQL streaming read replicas Aug 17, 2026
@O1eq
O1eq force-pushed the feat-add-treplica branch from d1b7036 to 2acbb24 Compare August 17, 2026 03:33
@O1eq O1eq changed the title Add treplica helper for PostgreSQL streaming read replicas feat: Add treplica helper for PostgreSQL streaming read replicas Aug 17, 2026
@O1eq
O1eq force-pushed the feat-add-treplica branch from 2acbb24 to 5eb79fb Compare August 27, 2026 23:15
Oleg Demeshev and others added 3 commits October 9, 2026 09:33

@codyfinegan codyfinegan left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Problems

  1. pgsql18 breaks. The volume is always mounted at /var/lib/postgresql/data (lines 248 and 274). pgsql18 uses PGDATA=/var/lib/postgresql/docker/pgdata with its volume at /var/lib/postgresql/docker (compose/pgsql.yml:226-232). The base backup ends up in the throwaway container and is lost. The standby then starts with an empty PGDATA, and the entrypoint tries to run initdb. The fix is to take the mount destination from docker inspect of the primary's .Mounts instead of hardcoding it.
  2. --port with no value loops forever. shift 2 with one argument left does nothing and returns an error, so while [[ $# -gt 0 ]] never ends. I reproduced this with the same parsing loop. A check like [[ -n "$2" ]] || { echo ...; exit 1; } fixes it.
  3. Restarting a stopped replica has a race. That branch runs docker start, then sleep 2, then do_status. If Postgres is not accepting connections yet, psql_replica returns nothing and the script prints "NOT in recovery mode" and exits 1. It should use the same wait loop as a fresh up.
  4. tdown probably errors while the replica is running. The replica container is not managed by compose but is attached to the totara network. docker compose down will then fail to remove the network ("active endpoints"). Someone should test this. The help text or tdown should mention it.
  5. A slot is left behind when the base backup fails. With --slot, the slot is created before pg_basebackup. The failure path removes the volume but not the slot, so the primary keeps WAL until someone runs down.
  6. Lag can read "in sync" when replication is broken. pg_last_wal_receive_lsn() keeps its last value after the WAL receiver disconnects, so receive LSN and replay LSN stay equal. Checking pg_stat_wal_receiver.status = 'streaming' would be more honest. The primary-side pg_stat_replication output shows the real state, so this is minor.
  7. Old versions are not supported, despite the "any pgsql version" claim. pgsql93 and pgsql96 default to wal_level=minimal and max_wal_senders=0, so pg_basebackup fails. The 9.x versions also lack pg_last_wal_*, replay_lag and, on 9.3, slots. It is enough to say PG10+ in the help.
  8. The down prompt runs before the script checks the primary. If the primary is stopped, the slot is not dropped, and the script prints nothing about it. It should warn when it skips this.

Smaller notes:

  • The added pg_hba line stays on the primary after down. That is harmless here because the primary is already trust, but down could say so.
  • max_wal_senders is also a setting where the standby must be at least equal to the primary. It matches by default, but it should go in the mirrored list.

The real blockers are items 1 and 2, and item 3 should be fixed too. The rest are follow-ups.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants