On this pageThe four things to keep separateExecution history is not the delivery queueOne queue has several identitiesSource trail: read in this orderTrace the periodic checkThe takeover that is easy to missQuestions to apply to any ownership proposalPractice before the discussionCheck yourself

Matching ownership: from routing to write fencing

This reading path uses public upstream Temporal code at 4db93cd73e. It explains the existing mechanisms that an ownership abstraction must preserve. It does not describe a shipped CDS implementation or reproduce an internal proposal.

The four things to keep separate

Mechanism Question it answers What it cannot establish by itself
Routing and membership Which host should receive this partition's requests? That an old host can no longer write
Persistence fencing Does this write still carry the expected ownership generation? That an idle loaded manager notices an intervening takeover promptly
Periodic validation Has this loaded queue lost ownership since it was acquired? Permission for a later write; ownership can change after the check
Task-ID allocation Which ordered IDs may this physical queue's writer allocate? Routing or ownership of another queue

Start with Task Queue, RangeID, and fencing token. History shard ownership and Matching queue ownership both use fencing concepts, but they protect different resources; do not assume they share the same counter or lifecycle.

Execution history is not the delivery queue

The cross-system comparison describes different storage architectures. Its “database-first” description of Temporal does not mean Temporal lacks an Event History: History records workflow events for replay, while Matching persists delivery tasks. A database or storage backend's write-ahead log is yet another layer. Changing the backing store for Matching does not, by itself, move workflow execution state out of History or change SDK replay semantics.

Keep three questions separate when discussing a storage migration: what data is stored, what makes that data durable, and what prevents a stale owner from changing it. The import graph answers none of these by itself; the source trail and behavior tests supply the evidence.

One queue has several identities

A namespace and queue family contain workflow/activity typed queues, partitions, and physical queues for versioned or unversioned work. Read the queue-family lab and upstream task_queue_id.go.

When reviewing a storage abstraction, write down which identity each operation consumes. A persistence name, a routing key, and an ownership token have different purposes even if the current implementation derives several of them from the same queue.

Source trail: read in this order

The Matching diagram shows package dependencies. Most of the following files belong to the same Go package, so they cannot appear as separate import nodes. Use this source trail for behavior inside that package.

Stop Read Look for
1. Routing matching_engine.go: watchMembership How routing changes lead to delayed partition unload
2. Periodic validation db.go: SyncState and verifyOwnershipLocked Dirty metadata / TTL refresh writes versus the read-only GetTaskQueue range check
3. Failure response pri_backlog_manager.go: signalIfFatal ConditionFailedError causing unload; periodic sync's error path
4. ID allocation backlog_manager.go: rangeIDToTaskIDBlock, task_writer.go Why changing a range counter also changes available task IDs
5. Store boundary task_manager.go, task_queues.go Manager/store translation and conditional queue updates
6. Regression evidence TestSyncState_UnloadsOnOwnershipLoss A simulated intervening owner increments the stored range; periodic sync must unload

Explore persistence packages. Read the legacy, priority, and fair backlog paths separately before assuming that a change to one covers them all.

Trace the periodic check

flowchart TD
    A[Loaded queue: periodic SyncState] --> B{Metadata changed or TTL refresh due?}
    B -->|Yes| C[Conditional UpdateTaskQueue]
    B -->|No| D[GetTaskQueue and compare stored RangeID]
    C --> E{Ownership conflict?}
    D --> E
    E -->|Yes| F[Existing error handling unloads the queue]
    E -->|No| G[Continue; operational errors follow their error path]

This is a behavioral reading aid, not a generated call graph. A successful read at time T does not authorize a write at T+1. The write must enforce its own condition.

The takeover that is easy to miss

Time Host A Host B Durable ownership
T0 Loads queue at range 10; remembers its local read/write state — Range 10
T1 Paused or temporarily misrouted Acquires queue, advances range, can write tasks Range 11
T2 Routing points back to A; its old manager still remembers 10 Stops or loses routing The intervening generation still matters
T3 Performs periodic validation or a conditional write with 10 — Mismatch exposes lost continuity; the old manager must unload

“Am I the routed host now?” cannot detect this history. A replacement ownership mechanism must also distinguish A's old ownership instance from a new one, even if both run in the same process. A read-only queue matters too: its in-memory read limits may not cover the intervening owner's task IDs.

Questions to apply to any ownership proposal

  1. Identity: What exactly is owned: a physical queue, partition, or a storage grouping? Which distinctions affect routing versus durable queue identity?
  2. Continuity: How is A → B → A detected? Is the expected generation attached to the loaded manager, or can it silently obtain a newer token?
  3. Write fence: Where is the ownership check atomic with the write? Which metadata, task, acknowledgement, and delete operations require it?
  4. Read-only validation: What is the maximum stale interval? What differentiates a definite loss from a timeout or unavailable backend?
  5. Task IDs: If allocation and ownership generations become independent, how are disjoint task-ID blocks and correct read limits preserved on takeover?
  6. Wrappers: Do retry, metrics, and rate-limit decorators preserve the extension's semantics? Do retries preserve the caller's expected generation?
  7. Lifecycle: Which loaded queues unload after loss? What prevents an in-flight initialization or callback from reviving stale state?
  8. Rollout: Which behavior is in scope now, and which migration or mixed-version questions belong to a separate plan?

These are review prompts derived from the existing invariants, not statements about an internal design's answers.

Practice before the discussion

Exercise What it teaches Limit
pnpm tour and pnpm wire Execution state versus delivery; RPC boundaries Teaching protocol/model, not a complete server
pnpm matching:lifecycle Loaded manager generation, namespace change, unload and reload Namespace failover is not a storage ownership transfer
Read Matching persistence and the queue persistence tests Queue range fencing, acknowledgement and recovery No CDS backend; local implementation differs from upstream
Read the upstream takeover test linked above Periodic detection of an intervening owner Does not establish the behavior of a future ownership checker
pnpm matching:failover-metrics Why manager lifetime affects metric label correctness Metric cleanup is separate from persistence fencing

Commands require a local tiny-temporal checkout. Check the alignment audit before generalizing a lab result to production.

Check yourself

Try answering before expanding:

Why does an idle queue need ownership validation?

It may not issue writes that would reveal a stale range. An intervening owner can change durable metadata and write task IDs beyond the old manager's view. Periodic validation detects lost continuity and causes the manager to unload.

Can an ownership cache replace write fencing?

No. Ownership can change immediately after the cached check. A write needs its own authoritative condition at the persistence boundary. A cached validation can only inform lifecycle decisions with an explicitly bounded delay.

Why is separating a lease token from RangeID more than renaming a field?

RangeID also drives task-ID block allocation in the current code. A separate ownership token must preserve queue-local allocation and read limits while binding every write and validation to the correct ownership instance.