DocsDemoGitHub
§4.4

The core model

Twelve tables: three vertex, four extracted edge, five derived edge. Structural only — everything here falls out of the CST plus the link pass.

The shape, in one pictureFile ──defines──▶ Occurrence ──references_local / resolves_to / calls │ ▲ / implements / type_defines──▶ Occurrence │ │ │ Scope ──contains──▶ Scope | Occurrence └──imports──▶ File

Two invariants

Everything below is shaped by two rules that later milestones lean on heavily.

Base rows are file-owned. Every base row carries the id of exactly one owning file, and nothing about one file's rows may depend on another file's. That is what lets the reduce phase rewrite a single file with a DELETE … WHERE file_id = $1 followed by a bulk copy, without reading anything else.

Identity is a string. A symbol is named by a SCIP-style descriptor, and cross-file resolution is a string match on it rather than an opaque id lookup — so the "symbol index" the link pass needs is just a btree on occurrence.descriptor. Seesymbol identity.

Vertex tables

Extracted, file-owned.

file
id, path, lang, pkg_scheme, pkg_manager, pkg_name, pkg_version
The unit of extraction, of ownership and of incrementality. The four pkg_ columns are the descriptor prefix — a package coordinate the batch supplies from go.mod (or package.json, …), never the file itself.
occurrence
id, file_id, descriptor, role, symbol_kind, name, range_start, range_end, scope_id
One place in one file where a symbol is defined or used. Definitions and references are the same row shape distinguished by role — which is what lets the link pass be a single self-join on descriptor.
scope
id, file_id, kind, range_start, range_end, parent_scope_id
The lexical containment skeleton. A scope says where a name could resolve, never what it resolves to.
§4.4 names file.file_id as the file table's primary key. gopgql requires every vertex type to declare a surrogate id: ID!, so the file's key is file.id and file_id survives as the owning-file column on the rows that are not files. Likewise, a gopgql edge table carries exactly (source_id, target_id) and no other column, so edge ownership is derived from an endpoint rather than stored: an intra-file edge is owned by the file its endpoints share, and a cross-file edge by its referencing file. Delete-by-file for edges is a join, not a column read.§4.4 gives contains one table with two possible endpoint types (scope → scope | occurrence). gopgql de-duplicates edge tables by name, and one table has exactly one source and one destination table, so a single contains could not carry both endpoint types. It is split into contains_scope and contains_occurrence, which share the one contains graph label — a MATCH on the label spans both, which is the §4.4 semantics, while each table keeps a real foreign key to its own target. That is why the schema is twelve tables rather than eleven.

Extracted edges — intra-file

Everything a file-local extractor can emit on its own, with no knowledge of any other file.

containsscope → scope | occurrenceLexical containment. Two physical tables — contains_scope and contains_occurrence — sharing one contains graph label, so each keeps a real foreign key to its own target while a MATCH on the label spans both.
definesfile → occurrence(definition)The file declares this definition.
references_localoccurrence(reference) → occurrence(definition)A same-file use, resolved during extraction because the target definition is in the same CST.

Derived edges — cross-file

None of these are extracted. Each is written by the link pass from base facts joined on descriptor, owned by the referencing file, and deleted and recomputed when that file changes. Materialising them is what makes cross-file navigation a read rather than a resolution at query time.

resolves_tooccurrence(reference) → occurrence(definition)A cross-file use, matched by descriptor.
importsfile → fileImport edges, by module or file descriptor. One Go import names a package, and a package is several files, so it becomes an edge per file.
callsoccurrence(definition) → occurrence(definition)The structural call graph. Approximate by construction — syntax plus a descriptor match, with no type resolution — and refined later by an overlay, never by reshaping the edge.
implementsoccurrence(definition) → occurrence(definition)The SCIP implementation relationship.
type_definesoccurrence → occurrence(definition)The SCIP type-definition relationship: this occurrence’s type is defined there. The source is any occurrence, not only a definition.
There is no link pass yet, so deploy/seed/seed.sql writes the cross-file edges out by hand in the place §7 will later compute them. They are part of the model, so they are part of what M1 demonstrates. M2 replaces those rows with a real full rebuild.

The SDL, in outline

One annotated GraphQL document is authoritative for all of the above.@node pins the label and physical table name,@relationship declares an edge and which table backs it,@column maps a camelCase field to its snake_case column, and@index / @check put the constraints in the database rather than in a loader that would have to be trusted.

schema/codiq.graphql — excerpttype Occurrence @node(label: "occurrence", table: "occurrence") { id: ID! fileId: ID! @column(name: "file_id") @index(name: "occurrence_file_id_idx", using: "btree") # THE join key of the whole system. descriptor: String! @index(name: "occurrence_descriptor_idx", using: "btree") role: String! @check(expr: "role IN ('definition', 'reference')") symbolKind: String! @column(name: "symbol_kind") name: String! @index(name: "occurrence_name_idx", using: "btree") rangeStart: Int! @column(name: "range_start") rangeEnd: Int! @column(name: "range_end") scopeId: ID @column(name: "scope_id") # Extracted, intra-file. definedIn: [File!]! @relationship(type: "defines", direction: IN) @hasInverse(field: "defines") referencesLocal: [Occurrence!]! @relationship(type: "references_local", direction: OUT) # Derived, cross-file — materialised by the link pass. resolvesTo: [Occurrence!]! @relationship(type: "resolves_to", direction: OUT) calls: [Occurrence!]! @relationship(type: "calls", direction: OUT) calledBy: [Occurrence!]! @relationship(type: "calls", direction: IN) @hasInverse(field: "calls") implements: [Occurrence!]! @relationship(type: "implements", direction: OUT) typeDefines: [Occurrence!]! @relationship(type: "type_defines", direction: OUT) }

Because the root field is derived from the table name rather than a pluralised label, queries read { occurrence … },{ file … } and { scope … }. Seethe query surface, orthe demo for worked exchanges against the seed.