The core model
Twelve tables: three vertex, four extracted edge, five derived edge. Structural only — everything here falls out of the CST plus the link pass.
The shape, in one pictureTwo invariants
Everything below is shaped by two rules that later milestones lean on heavily.
Base rows are file-owned. Every base row carries the id of exactly one owning file, and nothing about one file's rows may depend on another file's. That is what lets the reduce phase rewrite a single file with a DELETE … WHERE file_id = $1 followed by a bulk copy, without reading anything else.
Identity is a string. A symbol is named by a SCIP-style descriptor, and cross-file resolution is a string match on it rather than an opaque id lookup — so the "symbol index" the link pass needs is just a btree on occurrence.descriptor. Seesymbol identity.
Vertex tables
Extracted, file-owned.
fileid, path, lang, pkg_scheme, pkg_manager, pkg_name, pkg_versionoccurrenceid, file_id, descriptor, role, symbol_kind, name, range_start, range_end, scope_idscopeid, file_id, kind, range_start, range_end, parent_scope_idfile.file_id as the file table's primary key. gopgql requires every vertex type to declare a surrogate id: ID!, so the file's key is file.id and file_id survives as the owning-file column on the rows that are not files. Likewise, a gopgql edge table carries exactly (source_id, target_id) and no other column, so edge ownership is derived from an endpoint rather than stored: an intra-file edge is owned by the file its endpoints share, and a cross-file edge by its referencing file. Delete-by-file for edges is a join, not a column read.contains one table with two possible endpoint types (scope → scope | occurrence). gopgql de-duplicates edge tables by name, and one table has exactly one source and one destination table, so a single contains could not carry both endpoint types. It is split into contains_scope and contains_occurrence, which share the one contains graph label — a MATCH on the label spans both, which is the §4.4 semantics, while each table keeps a real foreign key to its own target. That is why the schema is twelve tables rather than eleven.Extracted edges — intra-file
Everything a file-local extractor can emit on its own, with no knowledge of any other file.
containsscope → scope | occurrenceLexical containment. Two physical tables — contains_scope and contains_occurrence — sharing one contains graph label, so each keeps a real foreign key to its own target while a MATCH on the label spans both.definesfile → occurrence(definition)The file declares this definition.references_localoccurrence(reference) → occurrence(definition)A same-file use, resolved during extraction because the target definition is in the same CST.Derived edges — cross-file
None of these are extracted. Each is written by the link pass from base facts joined on descriptor, owned by the referencing file, and deleted and recomputed when that file changes. Materialising them is what makes cross-file navigation a read rather than a resolution at query time.
resolves_tooccurrence(reference) → occurrence(definition)A cross-file use, matched by descriptor.importsfile → fileImport edges, by module or file descriptor. One Go import names a package, and a package is several files, so it becomes an edge per file.callsoccurrence(definition) → occurrence(definition)The structural call graph. Approximate by construction — syntax plus a descriptor match, with no type resolution — and refined later by an overlay, never by reshaping the edge.implementsoccurrence(definition) → occurrence(definition)The SCIP implementation relationship.type_definesoccurrence → occurrence(definition)The SCIP type-definition relationship: this occurrence’s type is defined there. The source is any occurrence, not only a definition.deploy/seed/seed.sql writes the cross-file edges out by hand in the place §7 will later compute them. They are part of the model, so they are part of what M1 demonstrates. M2 replaces those rows with a real full rebuild.The SDL, in outline
One annotated GraphQL document is authoritative for all of the above.@node pins the label and physical table name,@relationship declares an edge and which table backs it,@column maps a camelCase field to its snake_case column, and@index / @check put the constraints in the database rather than in a loader that would have to be trusted.
Because the root field is derived from the table name rather than a pluralised label, queries read { occurrence … },{ file … } and { scope … }. Seethe query surface, orthe demo for worked exchanges against the seed.