On this page1. Dispatch has two identities2. A queue name is a family3. OSSM-270: namespace change and unload4. Waiting polls, backlog, and propagation5. Proposed user-data storage migration6. Scheduling, versioning, and scaling policies7. Observability exerciseTeam workstreams to trace next

Matching lifecycle lab

Start with pnpm tour, then run pnpm matching:dispatch, pnpm matching:lifecycle, pnpm matching:user-data, and pnpm matching:policies. The six-step tour remains the shortest client-to-persistence path. Each new walkthrough prints queue identity, host and manager generation, namespace revision, durable backlog, loaded queues, waiting polls, user-data version, and metric labels.

1. Dispatch has two identities

QueueTask.scheduledEventId identifies the logical History event. A DeliveryRecord.taskId identifies one Matching backlog record within that physical queue. The task ID is ordered within a physical queue. A duplicate delivery can have a new task ID and the same scheduled event ID. On poll, Matching calls RecordWorkflowTaskStarted or RecordActivityTaskStarted on the History host that owns the run's shard. A success, TaskAlreadyStarted, NotFound, or a retryable failure decides whether Matching delivers, drops, or keeps the record.

The transfer processor removes History's transfer intent only after Matching accepts it. started lets Matching acknowledge its delivery record; it does not mean user code or the activity has finished. Activity completion validates the scheduled event and attempt in History. Run pnpm test and read test/dispatch.test.ts for the four failure probes.

Queue ownership now uses persisted range fencing and acknowledgement levels; see Matching persistence. Task tokens remain a teaching contract. The pinned OSS dispatch handling has more outcomes and retry logic.

2. A queue name is a family

QueueIdentity has four levels:

namespace + family name
  → workflow or activity typed queue
    → normal or sticky partition + partition ID
      → physical queue for a deployment version (or unversioned)

Database.taskQueues stores durable backlog by physical queue key. Partitions holds loaded partition managers and their physical-queue backlog managers in memory. Database.taskQueueMetadata stores each physical queue's durable range ID, acknowledgement level, and task-ID counter. A manager records the host that loaded it, manager generation, namespace revision, ownership generation, backlog gauge, and state. A membership change unloads the partitions now assigned to another host; test/cluster.test.ts moves a partition between two running Matching hosts. Normal startup uses one Matching host. The modulo hash is a small teaching substitute for upstream's consistent-hash ring. Compare the pinned OSS queue identities.

3. OSSM-270: namespace change and unload

The matching:lifecycle output prints identity, host, generation, namespace revision, durable backlog, loaded queues, waiting polls, in-flight work, user-data version, and metric labels at each step:

  1. Load a partition with an active namespace and one durable task.
  2. Change the namespace's active cluster. The registry callback stops and unloads managers and their user-data caches; durable backlog remains.
  3. Run an old gauge callback. Its manager is stopped, so the gauge stays zero.
  4. A passive poll receives no task. The toy still permits backlog storage and manager loading while passive.
  5. Change the namespace back and reload. A new manager generation reads the old backlog.
  6. Poll and complete the workflow.

Run pnpm matching:failover-metrics to see why the unload matters. Every backlog sample a partition emits carries a namespace_state tag that the engine fixes when the partition loads (loggerAndMetricsForPartition); the toy's Metrics handler and PartitionManager.namespaceState model that. The lab loads a partition while this cluster is standby, so its backlog reports as passive, then fails the namespace over here. With matchingUnloadOnNamespaceStateChange off, which is how upstream behaved before #12291, the partition stays loaded and every emitter tick keeps reporting the now-active backlog as passive, so a dashboard of active backlog reads zero on the active cluster until the partition idles out. With it on, the failover callback stops the partition, Stop zeroes its series under the old tags, and the next poll or add reloads it with active. test/metrics.test.ts covers both runs, a failover between two other clusters that leaves a passive partition loaded, and a delayed emit after unload.

test/lifecycle.test.ts also interleaves failover during initialization, host movement, delayed callbacks, and repeated stop. Namespace revision and ownership generation are independent. The pinned physical queue Stop path motivates stopping producers before zeroing backlog gauges. The team thread also calls out a stopping-partition race and its dependent cleanup PR. This lab captures the invariant, not those patches' exact code.

4. Waiting polls, backlog, and propagation

Matching.poll records worker, deadline, partition/version, and manager generation. An arriving task can sync-match a waiting poll without a backlog write. Otherwise it is stored; a later poll reads and acknowledges it. Cancellation, deadline expiry, and unload end waiting polls. Child adds try sync matching locally, then offer the task to their parent over AddWorkflowTask. Failed forwarding leaves backlog at the child; a reader retries ancestor delivery. A child poll is forwarded over PollWorkflowTaskQueue and waits at the root; cancellation crosses every hop. Sticky queues are rejected. See the alignment findings. test/polling.test.ts exercises each path and shows an Activity held in flight while an unrelated Workflow completes. See the polling and lifecycle walkthrough for the ownership sequence.

UserData keeps durable per-family configuration at the root workflow partition and propagates cached versions to child partitions only when asked. Delayed propagation is observable. Ephemeral statistics are separate from the durable snapshot. Root writes require an expected version, checked both in the manager and atomically in persistence. Two independent managers cannot overwrite the same version: one receives UserDataConflict. This follows the pinned OSS user-data topology.

5. Proposed user-data storage migration

The UserData migration flags are an exercise based on a draft design, not current OSS behavior. The legacy and CHASM rows live in the same TaskStore but behind separate read/write paths, staged with configureMigration. Try the sequence in test/user-data.test.ts: legacy read → seed and dual-write → shadow compare → CHASM read → completeness sweep → disable legacy writes → rollback. Keep this separate from a backlog-store migration. The draft CHASM design retains root-to-child propagation while changing the root's durable backing store. It flags cross-cell replication and in-memory component support as design considerations. The toy uses one global read gate and sequential writes to the two stores. Each write checks its expected version, but dual writes are not atomic and partial failures need reconciliation. It does not reproduce per-queue rollout, CHASM replication, or versioned transitions.

6. Scheduling, versioning, and scaling policies

dispatch-policy.ts uses priority first, then a simple served-count/weight score among fairness keys. Queue and per-key rate limits use a one-second clock window, and the limits themselves come from task-queue user data (queueRateLimit, perKeyRateLimits), so configuring dispatch is a root user-data write. Both sync match and backlog poll consult the policy, and tasks carry their fairness metadata from start-time scheduling options. This is a partition-local teaching approximation: fairness across partitions is not guaranteed, and the OSS algorithms are more involved. test/policy.test.ts shows sustained backlog and a limit that makes a waiting poll ineffective.

History's Execution.versioning supplies a directive on each transfer. Matching spools pinned tasks in their version's physical queue and auto tasks in default storage. At dispatch, current root user data selects the deployment for auto work; child/activity caches can lag until propagation. test/versioning.test.ts changes the deployment after tasks are already backlogged. This preserves late selection without moving backlog on every deployment change. The lab does not implement SDK deployment APIs or invalidation of already-spooled tasks when changing an execution's pinned behavior.

The scaling sketches live together in examples/src/scaling.ts. decidePartitions illustrates a server-side decision and tunePollers illustrates a separate worker-side decision, using synthetic backlog and poll observations. Neither is wired into Matching or an SDK worker, and their shared fixture is not an upstream feedback API. The example demonstrates shadow decisions and retaining backlogged partitions on scale-down. It does not implement autoscaling. See test/scaling.test.ts.

7. Observability exercise

Run pnpm matching:metrics. It prints two synthetic backlog samples with different partition and version labels, then illustrates an export that drops both dimensions. The collapsed total says a queue has backlog but cannot identify which partition or deployment needs polls. Keep namespace, queue family, type, partition, physical version, host, and manager generation available when investigating unloads or routing imbalance. Gauge cleanup after unload should be checked per complete label set: pnpm matching:failover-metrics shows a tag set that is correct at load and wrong after a failover, and inspect now prints emittedBacklogByNamespaceState next to the manager's own gauge.

Team workstreams to trace next

Workstream Start in this toy Status
OSSM-270 and stop/unload ownership matching:lifecycle, test/lifecycle.test.ts Runnable approximation
Dynamic partitioning examples/src/scaling.ts, child forwarding tests Synthetic policy sketch; no real partition fan-out
SDK poller tuning examples/src/scaling.ts, test/scaling.test.ts Synthetic example; worker poll counts are not tuned
Fairness service/matching/src/dispatch-policy.ts, test/policy.test.ts Partition-local approximation
Worker versioning version routing in service/matching/src/engine.ts, test/versioning.test.ts Pinned/auto routing example
Storage and user data common/persistence/memory/src/store.ts, test/user-data.test.ts Narrow interfaces; CHASM path is proposed design
Invalidation test/dispatch.test.ts Obsolete and duplicate task behavior
CDS recovery and task regeneration pnpm recovery PostgreSQL process restart recovery; independent store recovery is a future exercise

Worker control queues, heartbeats, eager dispatch, MCN, and Task Platform concepts remain exploratory extensions.