chlift preview

1 · The situation

Your app runs on Postgres. The daily-active-users chart now takes 5 seconds, and the weekly report times out. So you copy the data into ClickHouse.

On our own test data, a default Postgres took 5.4 s for "daily active users, last 30 days" and 22 s for "events per kind per week". The same queries took 0.1 to 1.6 seconds on ClickHouse. So the copy is worth doing. This page is about doing it without losing data.

2 · What goes wrong

The copy looks right. Then it quietly drifts.

We set up a table of 20,000 rows and replicated it to two ClickHouse replicas with Postgres’s default settings. Then, in Postgres, we updated 4,000 rows and deleted 1,000.

  • Deleted rows are still in ClickHouse. Postgres has 19,000 rows. Both replicas have 20,000.
  • Long text arrives empty. In one range of ids, the text column holds 124,215 bytes in Postgres and 1,335 bytes in ClickHouse.
  • Nothing reports an error. Replication shows as running. The only thing that noticed was a check that compared both sides.
The recording (3 min 21 s in real time; waits over 2 s are shortened here, so it plays in about 33 s). The safety setting is removed on purpose, the changes are made, and chlift migrate check fails. Text version: 07-corruption-repro.log.

3 · Why it happens

Postgres sends less than you think.

Logical replication sends each change with the table's replica identity. With the default identity, a delete carries only the primary key, and an update leaves out large values (TOASTed text and JSON) that did not change. ClickHouse tables are usually sorted by time first, so a delete that carries only the id cannot find the row it should remove, and the row stays. An update without the large value writes an empty one in its place.

4 · What to do (with or without chlift)

Three rules.

  1. Send the whole row. On every table that gets updates or deletes, before you start the copy:
    ALTER TABLE public.user_events REPLICA IDENTITY FULL;
    Postgres then writes the full old row to its WAL on each update and delete, so WAL grows. Plan disk space for it.
  2. Compare, don't count. A total row count hides errors that cancel out. Compare per day and per primary-key range: counts, ids, and checksums of the columns, on every replica.
  3. Cut over only on a match. Switch readers to ClickHouse only when every day and every key range is equal on both sides.

5 · How chlift does it

The same rules, as commands.

chlift is a command-line tool. It installs a ClickHouse cluster on your own Linux hosts over SSH (replicas, Keeper, and PeerDB for change data capture), then runs and proves the migration.

chlift install
Sets up the ClickHouse replicas, Keeper and PeerDB on your hosts. Running it again changes nothing unless the config changed.
chlift migrate plan
Picks the event tables and finds the risky ones (time-first sort keys, updated tables with long text or JSON). It writes the plan, including the REPLICA IDENTITY FULL lines, for you to review.
chlift migrate start --apply-postgres
Applies those Postgres settings, copies a snapshot to every replica, then keeps it current with live change data capture.
chlift migrate check --checksums
Compares row counts for every day and checksums for every primary-key range, on every replica.
chlift migrate cutover
Waits until replication has caught up, runs the check again, and stops replication only if 0 ranges differ. ClickHouse keeps the data.

With the setting in place, the same kind of changes gave 18,000 rows in Postgres and on both replicas, and 0 of 64 key ranges differed.

6 · What we measured

The full run, recorded.

One recorded run on 2026-10-03, from empty nodes and an empty ClickHouse. Every command is on a terminal recording with real timing, and nothing was typed or edited by hand. Setup: one shared 8-core Linux server (other workloads running too); two ClickHouse 26.3 replicas plus Keeper in containers capped at 1.5 GB RAM each; PeerDB on the same server; Postgres 16 with default settings in 768 MB; synthetic data from chlift seed (about 14M rows, 2.6 GB). This is not a vendor benchmark.

2m22s

Install from empty nodes

chlift install set up 3 hosts plus PeerDB, and all checks passed. The hosts are containers on one server; packages come from the official ClickHouse repo.

≤ 2m26s

14,000,000 rows into both replicas

8M + 4M + 2M rows copied by migrate start to both ClickHouse replicas. An upper bound: status was polled every 15 s.

0 of 64 buckets differ

Equal after live changes

After 50,000 inserts, 1,000 updates and 5,000 deletes over live CDC, migrate check --checksums found every day equal in Postgres and on both replicas (8,047,000 / 4,000,000 / 1,998,000 live rows). The checksums compare counts, ids, integers, timestamps and text lengths per primary-key range. JSON and float columns are not compared.

1m52s

Cutover, end to end

migrate cutover waits for PeerDB to confirm the watermark, then requires equal day counts and 0 of 64 differing checksum buckets on every replica for all 3 tables before it stops the mirror. ClickHouse keeps the data, and check still matches afterwards. Recorded with a later build that also compares rows below Postgres's lowest id.

Query time

14× to 163× faster than Postgres while PeerDB is still replicating (ClickHouse queries use FINAL). 31× to 308× faster without FINAL, which is what you get after cutover or for append-only tables.

Query (median of 5 runs, ms)PostgresClickHouse FINALClickHouse
Daily active users, last 30 days5,388266107
Events per kind per week, 6 months21,9501,600710
Top 10 accounts by events, last 7 days5,74015957
p95 latency per endpoint, last 30 days23,99414778
  • Median of 5 runs per query, all taken in one invocation. The raw samples and a script that recomputes this table are published with the recording.
  • Postgres time is measured from the client and includes sending the result; ClickHouse time is the server's own elapsed time.
  • Postgres runs with default settings in 768 MB. A tuned Postgres on bigger hardware would narrow the gap. Your data and queries are the real test.

7 · Install

One binary for Linux.

chlift v0.1.0-preview is one binary for Linux (amd64 or arm64). The script downloads it, checks it against the release's SHA256SUMS and against the checksums written into the script, and installs it only if both match.

curl -fsSL https://chlift.deemwar.com/install.sh | sh

Prefer to read it first: install.sh. Or download by hand: linux-amd64 · linux-arm64 · SHA256SUMS, then run sha256sum -c SHA256SUMS --ignore-missing.

Then

chlift init --cluster prod \
  --host a=10.0.0.1:clickhouse,keeper --host b=10.0.0.2:clickhouse,keeper --host c=10.0.0.3:keeper,peerdb
chlift plan                       # probes every host; changes nothing
chlift install                    # ClickHouse replicas, Keeper, PeerDB over SSH
chlift verify

export CHLIFT_PG_DSN='postgres://USER:YOUR_PASSWORD_HERE@db:5432/app'
chlift migrate plan               # writes chlift-migration.yaml for you to review
chlift migrate start --apply-postgres
chlift migrate check --checksums  # every day and every key range, on every replica
chlift migrate cutover
  • Hosts: Debian/Ubuntu (tested) or RHEL-family (written, not yet tested), SSH key access, Docker with compose v2 on the PeerDB host.
  • Postgres 12+ with wal_level=logical. Managed Postgres (RDS/Aurora) is not tested yet.
  • Preview: replicated targets stay preview until the fault soak passes.

What doesn’t work yet

Recorded end to end

Install from empty nodes; migrate plan, start, status, check with checksums; live CDC; cutover; the bench. The recording also turned up four bugs, all fixed in this release.

Still preview

A replication reliability run under injected faults is still in progress. Until it passes, replicated targets stay marked preview. We have not reproduced the known PeerDB row-loss issue with a Distributed target; chlift avoids that layout.

Help with it

chlift is free and MIT licensed. deemwar installs it, supports it and trains teams on it, for a fee. If you want that, tell us about your Postgres. A person replies, usually within a working day. Or email io@deemwar.com.

Running a quick security check…