DocsDemoGitHub

CodiQ

A distributed code-intelligence platform. Source code becomes a queryable property graph in PostgreSQL — and agents query it over MCP.

M1 — data model + MCPPostgreSQL 19 · SQL/PGQGraphQL over MCPPure Go, no cgo
Read the docsSee the demoGitHub

What an agent actually asks

Navigation questions — where is this defined, who calls it,what would a change here disturb — are traversals, not text searches. CodiQ stores them as edges so they are answered by a join instead of by re-reading the repository. The vocabulary is language-neutral, so one query spans a mixed-language codebase.

A question grep cannot answer, asked over the graph{ occurrence(name: "Put", role: "definition") { name descriptor definedIn { path } calledBy { name definedIn { path } } } }

That query, and four more, are answered over real seeded data onthe demo page.

Four principles

The model is the single source of truth

One annotated GraphQL SDL is authoritative. Every downstream artifact is generated from it — goose migrations, protobuf fact types, sqlc models and queries, and (via gopgql) the SQL/PGQ property-graph view and the GraphQL/MCP surface. No consumer is exempt.

Read/write separation

gopgql owns reads: GraphQL compiled to SQL/PGQ, served over MCP. A generated loader owns writes. The two never mix, and writes are never agent-triggered — the graph an agent sees is a read-only view over plain vertex and edge tables.

Dumb extractor, smart schema

The extractor emits file-local structural facts and nothing else. All derivation, cross-file resolution and transitive reasoning live behind the schema and query layer, where they can be recomputed rather than re-parsed.

File-disjoint incrementality

Base facts are file-owned: nothing about one file’s rows depends on another file. Re-indexing a file is a delete-by-file plus a bulk copy. Cross-file structure is derived by a link pass, not extracted.

The pipeline

A trigger starts a per-batch workflow. The map phase enqueues one extract task per changed file on a durable, Postgres-backed queue; the reduce phase runs once over the successful subset, in a single transaction. A failing file is retried in isolation, and a poison file is flagged and skipped rather than blocking the batch.

1

Extract

One task per changed file. A tree-sitter query plus a Go mapper turn the CST into occurrences, scopes and intra-file edges — pure Go, no cgo.

2

Transform

Facts are serialized to a protobuf artifact whose message types are generated from the SDL. The task checkpoints the artifact key, never the blob.

3

Load

One reduce step per batch, one transaction: delete-by-file, then CopyFrom straight into the target tables. No staging, no merge, no ON CONFLICT.

4

Link

Cross-file edges are materialized after load by matching reference descriptors against definition descriptors — a btree join, incrementally re-run over the affected neighborhood.

M1 ships the model and the read surface over hand-written seed data — there is no ingestion pipeline yet. The four stages above land progressively from M2 onwards, each as a working, integration-tested slice rather than a shim. See the pipeline docs for what is built when.

What M1 ships

The data model is the product in M1. The SDL defines the core vocabulary; gopgql — pulled as a container image — generates the full schema into postgres:19beta2, creates the property-graph view over it, and serves GraphQL/MCP. Rows are inserted directly from a seed file, because no extractor exists yet.

The core model →

file, occurrence and scope, the intra-file edges extracted with them, and the cross-file edges derived afterwards.

Symbol identity →

Every symbol is named by a SCIP-style descriptor. Identity is a single human-readable string; resolution is a string match, not an opaque ID.

Run it locally →

Docker Compose brings up Postgres 19 and gopgql, applies the generated migrations and loads the seed. One command.