DocsDemoGitHub
§5–§9

The ingestion pipeline

extract → transform → load → link. This is the end state; M1 ships none of it, and the milestones converge on it by optimising an already-working slice.

M1 has no ingestion at all — rows are inserted from a seed file. Everything on this page lands from M2 onwards, and each milestone below is a working, integration-tested slice rather than a step towards a big-bang cutover. The badge on each section says when it arrives.

Shape

trigger (push / manual / watch) │ ▼ batch workflow ──▶ MAP — per file, durable queue │ extract A ──▶ artifact A │ extract B ──▶ artifact B │ extract N ──▶ artifact N ▼ REDUCE — one transaction, over the successful subset delete-by-file + CopyFrom ──▶ incremental re-link │ ▼ PostgreSQL 19 (vertex / edge tables) ──▶ gopgql ──▶ agents

Extract M2

The engine is gotreesitter, a pure-Go tree-sitter runtime: no cgo, so the binary cross-compiles to any GOOS/GOARCH, with an embedded grammar registry and one query engine spanning every language.

A language is authored as a .scm tree-sitter query plus a Go mapper. The query captures CST nodes using a shared capture vocabulary — this is where normalisation to the neutral core happens:

@definition.<kind>A definition occurrence — function, method, type, field, variable, module, …
@reference.<role>A reference occurrence — call, read, write, type.
@scope.<kind>A lexical scope.
@importAn import occurrence.
@nameThe identifier used for the descriptor and the name column.

The mapper builds the descriptor suffix from the capture hierarchy, assignssymbol_kind and role, resolves same-file references to local definitions, and emits the base rows. Anything requiring another file is emitted as an unresolved reference carrying its target descriptor. Depth is structural only; a go/types-style semantic producer can populate overlay tables later without changing this contract.

Adding a language touches no core schema:

+ extract/<ext>/<ext>.go # the mapper (captures → facts) + extract/<ext>/query.scm # the tree-sitter query, //go:embed'd + coord/<ecosystem>.go # the package-coordinate resolver ~ extract/extract.go # register ".<ext>" in byExt ~ coord/coord.go # register the resolver

Transform M5

Facts are serialized to a protobuf artifact whose message types are generated from the SDL — SDL to .proto, then buf to Go, with buf lint and breaking-change detection guarding schema evolution. The artifact goes to a shared volume, and the map task checkpoints the artifact key, never the blob.

Artifacts are short-lived: they survive until the batch's reduce completes, then they are deleted. A failed batch keeps its artifacts, so a retry consumes them without re-extracting anything.

Load M2batched at M4

One reduce step per batch, in a single transaction over the successful map subset. Per file: DELETE FROM <table> WHERE file_id = $1, then CopyFrom the file's new rows straight into thetarget tables.

No staging table, no merge, no ON CONFLICT — deleting first guarantees there is no key to collide with. That simplification is bought entirely by file-disjointness: because a file owns its rows outright, deletion is a single predicate.

Vertices load before edges within each file, and intra-file edge endpoints are co-loaded. The sequence inside the one transaction is: base load, core link, then any overlay producers.

Link M2incremental at M8

Cross-file edges are derived, written after load rather than extracted. The join key is the descriptor already sitting on base rows, and the index is the btree on it — there is no separate structure to build or invalidate.

A cross-file edge is owned by its referencing file: re-linking a file deletes its cross-file edges and recomputes them, and moving a definition triggers a re-link of the files that reference it, found through the same descriptor index. On change, only the affected neighbourhood is recomputed — the files referencing the changed file's definitions, plus the definitions the changed file references.

A full re-link — delete every cross-file edge, recompute from base facts — stays available as a backstop, run nightly on a schedule and on demand, so an incremental-invalidation bug self-heals rather than quietly accumulating drift.

Orchestration M3

DBOS Transact is an embedded library over the Postgres CodiQ already runs — no separate orchestration server, no extra infrastructure. Each stage is a checkpointed step, so recovery never re-runs completed work; a crash mid-batch resumes.

Per-task retry, timeout and flagging mean a failing file is retried in isolation and a poison file is skipped rather than blocking the batch. DBOS's own system tables live in a separate database(codiq_dbos) on the same instance, so checkpoint writes do not contend with the reduce phase's bulk copy.

Why this order

The pipeline is not built front-to-back. M2 is a single program with goroutines, no durability and no batching, and it works end to end. M3 makes it crash-resumable, M4 makes it a batched map-reduce, M5 moves the facts off the heap and onto disk as protobuf. Each step optimises something that was already shipping.

Back to what M1 ships, or on torunning it locally.