Skip to content

Transaction Engine

ParticleDB’s transaction behavior is controlled by two server-start choices:

  • --wal-sync-mode decides durability cost (sync, groupsync, nosync)
  • --txn-mode decides concurrency behavior (fast, occ, serializable)

The codebase no longer has a single always-on “serializable snapshot isolation” path. Instead, the engine exposes several operational modes so you can choose between maximum throughput and stronger safety.

The write-ahead log persists mutations before they become durable.

ModeDurable when COMMIT returns?Batching policy
syncYesfsync forced at the commit site
groupsync (default)YesConcurrent commits share one fsync
nosyncNono fsync

Both sync and groupsync are durable-on-return. They differ in when the fsync is issued, not in whether your commit is on disk before you are told it succeeded. groupsync amortizes one fsync across many concurrent committers, which is why it is the default: under load, most committers perform no disk I/O of their own.

Terminal window
particledb start --wal-sync-mode sync
particledb start --wal-sync-mode groupsync
particledb start --wal-sync-mode nosync

nosync performs no fsync at all. Data written since the last flush is lost on power failure. Use it only for benchmarks and reloadable data.

--txn-mode sets the server-wide default isolation level. Individual sessions may override it with SET TRANSACTION ISOLATION LEVEL (see below).

There are three supported modes.

Terminal window
particledb start --txn-mode fast # no isolation guarantee
particledb start --txn-mode occ # default
particledb start --txn-mode serializable
ModeIsolation levelAllowsApproach
fastNoneconcurrent writers overwriting each otherWrites apply in place. No commit-time validation, no conflict detection. Bulk load and single-writer ingest only.
occ (default)Snapshot Isolationpredicate anti-dependency cycles (G2)Optimistic: transactions buffer their writes and are validated at commit; reads see a consistent MVCC snapshot
serializableStrict SerializablenothingSerializable plus real-time ordering
Phenomenonoccserializable
G0 — dirty writepreventedprevented
G1a — aborted readpreventedprevented
G1b — intermediate readpreventedprevented
G1c — circular information flowpreventedprevented
P4 — lost updatepreventedprevented
A5A / G-single — read skewpreventedprevented
A5B / G2-item — write skewprevented²prevented
G2 — predicate anti-dependencyallowedprevented
Real-time — a transaction beginning after another’s acknowledged commit observes itnot guaranteed¹prevented

¹ Only serializable guarantees real-time ordering, and it does so by construction: serialized publish under the certifier lock plus commit-wait. occ has no such construction and therefore carries no such guarantee. In practice it does much better than that phrasing suggests — 12,000 probes across both modes and both durable WAL modes, on both wire protocols, with a control that detects the anomaly on 100% of iterations when a snapshot is genuinely pinned early, found none on a single node. Treat that as a measurement, not a promise: if you need the ordering, use serializable.

² Measured with the Hermitage harness. Textbook Snapshot Isolation permits write skew, but occ validates the rows a transaction actually read, and that row-level read set catches the item case. It cannot catch the predicate case — see below.

Two scope limits worth knowing before you rely on the occ measurement.

Single node only. We have not measured it on a replicated cluster.

--oltp-partitions > 0 breaks it. With partition workers enabled, an explicit BEGIN on the main port may not observe their commits. serializable refuses that flag for exactly this reason.

What occ still permits, and serializable does not, is the predicate anti-dependency cycle (G2). occ validates the rows a transaction actually read, so it cannot see a row that was never read because it did not yet exist. When a conclusion depends on the absence of rows matching a predicate — “no overlapping booking exists”, “at least one person is still on call” — a concurrent insert that would have matched is invisible to validation.

Practical rule: if your invariant is about rows you read, occ holds it. If it is about rows that must not exist, use serializable.

--txn-mode sets the server-wide default. A session may select its own level, and the requested level is applied to the engine — not merely recorded.

Level requestedEngine modeReported labelRelative to the occ default
READ UNCOMMITTEDRead Committed (promoted, as in PostgreSQL)read committedweaker
READ COMMITTEDRead Committed — fresh snapshot per statementread committedweaker
REPEATABLE READSnapshot Isolation — one snapshot for the transactionrepeatable readsame
SERIALIZABLESerializableserializablestronger

Note that READ COMMITTED — PostgreSQL’s default, and therefore something ORMs and pools often set explicitly as a no-op — is a downgrade here, because our default is Snapshot Isolation. REPEATABLE READ is the request that matches the default.

Anything that is not one of the four standard levels is rejected with 42601 (invalid transaction mode list). READ ONLY / READ WRITE are honored (last-one-wins, and writes inside READ ONLY fail with 25006); [NOT] DEFERRABLE is parsed and ignored.

If --txn-mode is omitted the mode is occ. A durability flag does not change your isolation level.

The one exception is --wal-sync-mode sync with --txn-mode omitted. That combination used to select a serializable engine through an inference that no longer exists, so silently defaulting it to occ would weaken an existing deployment to snapshot isolation. The server refuses to start and asks you to state the isolation level.

Earlier builds accepted several other spellings. They are removed, not deprecated — the server refuses to start and names the replacement.

RemovedUse insteadWhy
occ-v3, occv3, occ-v2, occv2, zero-copy-occ, zcocc, serializable-occoccAll were already names for the same engine
strict-serializable, strict-serserializableSame mode, one canonical name
serializable-ssi, ssiserializableStrict serializable is strictly stronger; no application correct under SSI can break on it
table-2pl, 2pl, row-2pl, row_2plserializableThese had already stopped being two-phase locking and mapped to the SSI policy

Moving from serializable-ssi to serializable is safe for correctness, but a high-concurrency, low-conflict write workload may be meaningfully slower. Measure before migrating.

Within an open transaction, mutations are buffered before commit:

  • INSERT, UPDATE, and DELETE are collected in a transaction buffer
  • reads merge the base snapshot with the transaction overlay
  • COMMIT replays buffered operations
  • ROLLBACK discards them

That overlay is what allows statements later in the same transaction to see rows inserted or updated earlier in the same transaction.

The core lifecycle is standard:

BEGIN;
UPDATE accounts SET balance = balance - 100 WHERE id = 1;
UPDATE accounts SET balance = balance + 100 WHERE id = 2;
COMMIT;

Two-phase commit is also implemented:

BEGIN;
UPDATE accounts SET balance = balance - 500 WHERE id = 1;
PREPARE TRANSACTION 'transfer_tx_001';
COMMIT PREPARED 'transfer_tx_001';
-- or
ROLLBACK PREPARED 'transfer_tx_001';

ParticleDB parses and executes the common PostgreSQL row-locking forms:

  • FOR UPDATE
  • FOR SHARE
  • NOWAIT
  • SKIP LOCKED

Example:

BEGIN;
SELECT *
FROM jobs
WHERE status = 'pending'
ORDER BY created_at
LIMIT 1
FOR UPDATE SKIP LOCKED;
COMMIT;

The finer PostgreSQL lock modes FOR NO KEY UPDATE and FOR KEY SHARE are not part of the current implementation.

In practice:

  • --txn-mode sets the server-wide default; sessions may override it
  • use normal BEGIN / COMMIT / ROLLBACK in clients
  • SET TRANSACTION ISOLATION LEVEL is honored — see the table above
  • SAVEPOINT, RELEASE SAVEPOINT, and ROLLBACK TO SAVEPOINT are supported
  • groupsync + occ is the default and the right starting point: durable commits at Snapshot Isolation.
  • Move to serializable when correctness depends on ordering the database cannot observe — for example when clients coordinate through a queue or an external service.
  • Use nosync + fast only when you can tolerate data loss on crash.
  • Use --oltp-partitions for warehouse-affine OLTP workloads such as TPC-C.
  • Keep transactions short so buffered writes and lock hold times stay small.