A plan declares its types; a consumer derives them again — or copies what
the plan declared. When the two disagree, nothing in the format notices. Each cell below is
one of 108 plans read by one of 9 implementations, compared against an
expectation written from the spec rather than from any implementation.
An empty cell is silence, not a wrong answer. It means the
implementation does not accept that plan at all — which is why silence is drawn as an
outline and never as a colour.
A second corpus is measured below the matrix: 71 relation cases
written by hand against the sentences of the relation documentation, read by
4 implementations — what they answer.
Click or tap a cell to open details below its row. The outlined cell is the
source; a GitHub mark means an issue or PR is linked.
What the grid cannot show
A row of agreement is worth what the case behind it pins. deriver/ derives each
schema a second time from the specification, and deriver/mutants.py then replaces
each of its rules with another reading of the same sentence and rederives: a reading no case
tells apart is a rule this corpus states and tests with nothing. The only check column
marks the 10 cases that are the only ones to pin some reading — remove one and that
rule goes unchecked.
4 readings are pinned by nothing:
read: the mask is read in the order its items are listed
set: the field types come from the last input
write: the output is the declared table schema
functions: a signature binds any impl of its function
Two of the three cannot be reached by a valid plan at all — the specification
requires a set operation's inputs to agree on types and a write's input to match its
table_schema, so no legal plan distinguishes the readings. The third is a question
the specification leaves open, and the deriver declines such a plan rather than picking a side.
The relation corpus
A second measurement, on a second corpus: 71 cases written by hand against
the sentences of the relation documentation, compiled to protobuf and read by 4
implementations. It shares no case with the matrix above, so it shares no column.
substrait-java58 matched (42 on the schema alone), 5 not accepted
DataFusion43 matched, 3 differed, 17 not accepted
substrait-go40 matched (36 on the schema alone), 2 differed, 21 not accepted
DuckDB27 matched, 11 differed, 25 not accepted
matchedmatched, rows not observeddifferednot acceptedobserved, never scored
47 of the 71 cases assert rows as well as a schema, and only DataFusion, DuckDB executes, so every other agreement on those is an agreement about half of what the case says. Two of them, a right semi join and a right anti join, emit the same columns with the same nullability and are separated by nothing but their rows.
Click or tap a cell to open what that participant answered, what the case
asserts, and where they differ, why. A GitHub mark means an issue or PR is linked.
The saved columns are
results/relations/, one file per
participant; probe/relations/replay.sh rebuilds one from nothing and requires it
back. The cases are tests/relations/, and the
reasons behind the differing cells are
differed.json.