Repository navigation
Conversation
NingZhou-NZ
reviewed
Aug 10, 2026
O1eq
force-pushed
the
feat-add-treplica
branch
from
August 17, 2026 03:33
d1b7036 to
2acbb24
Compare
O1eq
force-pushed
the
feat-add-treplica
branch
from
August 27, 2026 23:15
2acbb24 to
5eb79fb
Compare
…ics, not Mimir Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
codyfinegan
force-pushed
the
feat-add-treplica
branch
from
October 8, 2026 20:33
3acf042 to
daf56ef
Compare
codyfinegan
requested changes
Oct 8, 2026
codyfinegan
left a comment
Member
There was a problem hiding this comment.
Problems
- pgsql18 breaks. The volume is always mounted at /var/lib/postgresql/data (lines 248 and 274). pgsql18 uses PGDATA=/var/lib/postgresql/docker/pgdata with its volume at /var/lib/postgresql/docker (compose/pgsql.yml:226-232). The base backup ends up in the throwaway container and is lost. The standby then starts with an empty PGDATA, and the entrypoint tries to run initdb. The fix is to take the mount destination from docker inspect of the primary's .Mounts instead of hardcoding it.
- --port with no value loops forever. shift 2 with one argument left does nothing and returns an error, so while [[ $# -gt 0 ]] never ends. I reproduced this with the same parsing loop. A check like [[ -n "$2" ]] || { echo ...; exit 1; } fixes it.
- Restarting a stopped replica has a race. That branch runs docker start, then sleep 2, then do_status. If Postgres is not accepting connections yet, psql_replica returns nothing and the script prints "NOT in recovery mode" and exits 1. It should use the same wait loop as a fresh up.
- tdown probably errors while the replica is running. The replica container is not managed by compose but is attached to the totara network. docker compose down will then fail to remove the network ("active endpoints"). Someone should test this. The help text or tdown should mention it.
- A slot is left behind when the base backup fails. With --slot, the slot is created before pg_basebackup. The failure path removes the volume but not the slot, so the primary keeps WAL until someone runs down.
- Lag can read "in sync" when replication is broken. pg_last_wal_receive_lsn() keeps its last value after the WAL receiver disconnects, so receive LSN and replay LSN stay equal. Checking pg_stat_wal_receiver.status = 'streaming' would be more honest. The primary-side pg_stat_replication output shows the real state, so this is minor.
- Old versions are not supported, despite the "any pgsql version" claim. pgsql93 and pgsql96 default to wal_level=minimal and max_wal_senders=0, so pg_basebackup fails. The 9.x versions also lack pg_last_wal_*, replay_lag and, on 9.3, slots. It is enough to say PG10+ in the help.
- The down prompt runs before the script checks the primary. If the primary is stopped, the slot is not dropped, and the script prints nothing about it. It should warn when it skips this.
Smaller notes:
- The added pg_hba line stays on the primary after down. That is harmless here because the primary is already trust, but down could say so.
- max_wal_senders is also a setting where the standby must be at least equal to the primary. It matches by default, but it should go in the mirrored list.
The real blockers are items 1 and 2, and item 3 should be fixed too. The rest are follow-ups.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Adds a new
bin/treplicacommand that creates and manages a PostgreSQL streaming read replica of any totara pgsql container. Its purpose is testing Totara's database read/write splitting feature ($CFG->dboptions['readonly']), which routes read queries to a read-only replica.treplica up [pgsqlNN]- clones the running primary withpg_basebackupinto a new volume and starts it as a hot standby (default service: pgsql16). Image,PGDATAand resource settings (max_connections,max_locks_per_transaction, ...) are derived from the running primary, so it works with any pgsql version/config without hardcoded values. Idempotent: re-running shows status, a stopped replica is restarted, a stale volume is refused with a hint. Ends by printing the ready-to-pasteconfig.phpblock (including a warning that it must be placed after theconfig-after.phpinclude, which reassigns$CFG->dboptions).treplica status- recovery check,pg_stat_replication, replication lag.treplica logs- follows the standby's statement log.treplica down- removes container + volume after confirmation and drop the replication slot if one was created.--port(publish on a host port),--slot(physical replication slot for bulk-load safety),--quiet-logs.Testing Instructions
pgsql16)treplica up- it should finish withstate: streaming/Replica lag: in syncand print a config.php snippettreplica status- same healthy output; runtreplica upagain - it should no-op politelyreadonlyblock at the end of your config.php, set$CFG->perfdebug = 15;, reload any page and check the footer shows "DB reads from replica" counting up, whiletreplica logsshows only SELECT statementstreplica down, confirm - container, volume (and slot, if--slotwas used) are removed;treplica statusnow reports there is no replicaChecklist
bin/directory run correctly on both MacOS and WSLtbuild container && tup container)config.phpare compatible with our oldest supported Totara version, our newest Totara version, and Moodle