Changes & retractions

変更と撤回

出した結果を後から取り下げたときの記録を、まとめてここに置く。 サイトに載せていない業種も含める。載せているかどうかは読者の都合であって、 撤回があったかどうかとは関係がない。

いまの版

業種現行版版の記録DOI(この版)
通関 HS分類v1.0.0010.5281/zenodo.21847252
税務 根拠条文v1.1.0110.5281/zenodo.21873243
建築2D図面 DXFv1.1.0110.5281/zenodo.21896774
積算(未掲載)v1.0.0010.5281/zenodo.21847244
土木CAD SXF(未掲載)v1.1.2310.5281/zenodo.21896731
地盤 液状化判定(未掲載)v1.0.0010.5281/zenodo.21847240
機械 STEP AP242(未掲載)v1.2.0210.5281/zenodo.21887256
建築BIM IFC(未掲載)v1.2.1310.5281/zenodo.21896730
方法論論文(プレプリント)(未掲載)v1.3.1510.5281/zenodo.21896705

この表もこの下の原文も、各リポジトリの .zenodo.jsonCITATION.cff から生成している。このページに手で書いた文は無い。 撤回の記録を2箇所に持つと、片方が腐るため。

税務 根拠条文 zeimu-bench

Version 1.1.0 records a defect present since v1.0.0 and not previously disclosed. The benchmark asks which statutory provisions the tax authority relied on. In 11 of 227 graded citation elements (4.8%), the enquiry text itself names the provision that is the answer — sometimes in the title. This is not an authoring error: the enquiries are the authority's own published text, and a questioner citing a provision is ordinary practice. But the grader does not distinguish a provision that was researched from one copied out of the question, so scores are inflated by an undisclosed amount. A check is added and the count is calibrated against it.

建築2D図面 DXF cad-bench

Version 1.1.0 adds a fifth task, post-hoc baseline records for all five, and the repair of a defect this repository introduced into itself. A hash freeze was added to protect the five published reference drawings; because those drawings are generated from a specification and a DXF embeds a creation timestamp and fresh object identifiers, the run script regenerated them on every run and the commit that introduced the freeze rewrote all five files it was written to protect. The geometry and every score were identical, so no reported figure moved, but bytes released under a DOI were changed by the act of checking that they should not change. The rule adopted is to hash the inputs, not the outputs: a generated artefact is verified by regenerating it to a temporary file and comparing meaning — re-scoring old against new — rather than bytes. The run script no longer regenerates the reference at all, and creates it only when absent, which closed the root cause rather than the workaround of restoring the files by hand after each run. Testing that path exposed a further gap: with the reference deleted, the freeze verifier passed, so it now fails when a task has answers but no reference.

土木CAD SXF doboku-bench

Version 1.1.2 adds a single entry point and, in building it, exposed two holes in the checks. This was the only one of the eight repositories without a script that runs every check, so each had to be invoked by hand with the right arguments — and consequently nobody ever invoked them all. T001 and T002 carried no freeze or baseline record at all, because the post-hoc baseline mode was implemented in only one of the eight repositories. Separately, the function mapping a variant task name to its base directory split the name on the letter 'n', turning 'T002clean' into 'T002clea'; the verifier looked in a directory that does not exist, reported no record found, and returned success, so that variant's freeze record had never once been checked. A missing record is not drift, so a checker built to detect drift passes. Both are fixed and calibrated, and the run script no longer regenerates the reference file on each run. No score changed: arm A remains 26.0, 26.0 and 92.0, arms B and C remain 100, and the clean-material arm remains 100 in all three runs.

Version 1.1.1 resolves one of the two qualifications v1.1.0 introduced. v1.1.0 reported that the material supplied to the format-informed arm contains four rows of the reference solution, and that the 25-to-97 effect therefore could not be separated into 'explaining the format' and 'handing over answers'. It has now been separated. Rather than revise the material — revising after reading submissions is the failure that forced a withdrawal in a sibling benchmark — the leaking document was left in place and a second copy issued with every example value and entity number replaced by one that does not occur in the task, while the explanatory content (the standard colour, line-type and line-width tables, the ordering rule, the full-scale-coordinate rule) was left unchanged. Before running we verified that no example row matches the reference and that all five parse with the expected argument counts, so the document remains a working specimen of the format. The arm scored 100 in all three runs, exactly as with the leaking material. The four disclosed rows were not carrying the score, and the material effect can be read as the effect of explaining the format.

Version 1.1.0 withdraws two statements from v1.0.0. The unsupplied arm's failure does not replicate: run three times on the same task under the same condition it scored 26, 26 and 92, the whole spread turning on one binary guess — whether each record is wrapped in the format's comment delimiters or written as a numbered instance. All three runs named that guess among their own top risks before submitting; one ranked it first, called it fatal if wrong, and got it wrong. The v1.0.0 figure of 26.0 is therefore one draw from a coin, not a property of the arm. The accompanying observation that the arm 'ranked its risks and was killed by the third' is withdrawn: the ranking itself varies between runs.

機械 STEP AP242 kikai-bench

Version 1.2.0 completes a 2x2 that isolates the cause of the v1.1.0 defect. T001 scored 94.6 in all six runs because nine graded fields were returned null by every one of them, and two examination-side properties could each explain it: the answer format's worked example was structurally incapable of showing the field, and the question text never asked for it. Three successor tasks vary those two independently while sharing T001's reference byte for byte. T003 repairs the worked example and leaves the question untouched; T004 additionally rewrites the question to name all five graded fields; T005 removes the field from the example again — omitting the key rather than falsifying its value — while the question still asks for it. All three score 100.0, and the nine previously-null items match the reference exactly in every run. Either disclosure route alone is sufficient; T001 failed because neither was present.

Version 1.1.0 records three defects in the answer format's worked example, present since v1.0.0. The example is simultaneously specimen and specification, and it fails in three directions at once. Disclosure: its eight fields match the reference exactly, so one of twenty-eight graded tolerance records is handed over. Error: the datum list it shows is shorter than the file's actual list, and nothing in the task says the example is abridged, so it presents a value that is not the answer for a field that is graded. Concealment: six graded fields never appear with a value — the cause of the defect already reported in v1.0.0, where six independent runs returned the same nine nulls because the field's existence was never disclosed. The second task's example conceals eleven. The three pull against each other: avoiding disclosure by inventing an example produces a row that does not exist, and choosing a real row discloses it. The correct form is a row that is real but not graded. The example is left unrepaired and a check is added.

建築BIM IFC bim-bench

Version 1.2.1 adds the post-hoc baseline records that were missing from T001 through T005. The accompanying paper stated that already-run tasks carried a baseline hash record; in this repository five of them did not, and the freeze verifier passed anyway, because a missing record is not drift. The baseline mode itself had never been implemented here — only the pre-run freeze, which refuses to stamp once answers exist — so there was no command that could record these tasks at all. A baseline does not certify the past; it only makes later drift detectable, and the records name what could not be captured: the solver prompts for these five tasks were never retained.

Version 1.2.0 reports that repairing the v1.1.0 defect introduced a worse one. The note written to document it printed the answers to four of the repaired task's thirteen questions — the four the repair existed to ask. All three solver runs reported it independently; five checks passed the statement carrying them, including a scope check written for the previous defect. The task is recorded as a failed measurement and left unrepaired.

Version 1.1.0 withdraws the principal finding of v1.0.0, that the arm permitted to execute code led on every axis. The 35.7-point separation it rested on came from a failed measurement rather than a result: a submission written without opening the corpus scored 100.0 where all six real runs scored between 55.0 and 70.0, because the adversarial suite was generated by perturbing the reference and therefore never probed submissions far from it. The task statement and the grader had also been revised after the submissions were read, by 52.6 points — more than the separation being claimed, so the pilot and the replication were not runs of the same task. A replacement task, frozen before any arm ran, put both arms at 11 of 11 and separated them on cost alone.

方法論論文(プレプリント) ai-reach-paper

Version 1.3.1 corrects a statement in v1.3.0 that was false when published. The threats section said that tasks which had already run carried a post-hoc baseline hash record. In two of the eight repositories they did not: six tasks with answers carried no record of any kind, and a seventh carried one that was never read, because the function mapping a variant name to its base directory split the name on the letter 'n' and turned 'T002clean' into 'T002clea'. The verifier looked in a directory that does not exist, reported no record found, and returned success — a missing record is not drift, so a checker built to detect drift passes. Both followed from the same omission: the baseline mode existed in only one of the eight repositories, so elsewhere no command could record an already-run task at all. The records are now in place and the statement is true, dated rather than silently repaired.

Version 1.3.0 replaces an interpretation with a controlled experiment, and corrects two of the paper's own figures. The examination-side defect reported in v1.1.0 was diagnosed as the solvers answering what was asked; that was a causal claim and it had not been measured. Three successor tasks that share the reference byte for byte vary the two disclosure routes independently — whether the worked example shows the graded field, and whether the question text asks for it — completing a 2x2 with the original. Either route alone restores full agreement (100.0 in every cell but the original, which scored 94.6), so 5.4 points of a published score were a property of the examination rather than of the solvers. A solver identified the design's limit: the task naming the omitted key makes the condition 'asked, with the name available' rather than 'asked alone', and this is not removable, so the factor actually varied is whether the graded field's existence and name are discoverable at all. This does not soften the paper's central claim — the 2x2 was only constructible because replication had already exposed which field was affected. Controlled substitution can localise an examination-side defect; it cannot find one.

Version 1.2.1 resolves one of the two qualifications v1.2.0 placed on the series' largest material effect. v1.2.0 reported that the material supplied to the format-informed arm contained four rows of the reference solution, so that 'explaining the format' and 'handing over answers' were confounded. The confound has been measured rather than argued: the leaking document was left in place, a second copy was issued with every example value and entity number replaced and the explanatory content unchanged, and the arm scored 100 in all three runs, exactly as before. The disclosed rows were not carrying the score. Both limits on that inference were declared before the runs — the replacement changed two things rather than one, so only the absence of a drop is interpretable, and 100 is a ceiling — and both are stated in the text. The other qualification, that the unsupplied arm's failure does not replicate, is unchanged.

Version 1.2.0 reports a fourth and fifth kind of examination-side defect, both of which we introduced ourselves. Repairing the task-statement defect of v1.1.0 produced a note that printed the answers to four of the repaired task's thirteen questions; all three solver runs reported it, and five checks — calibration, external cross-check, adversarial suite, a new scope check written for the previous defect, and a hash freeze — passed the statement carrying them. Auditing all eight benchmarks for the same fault then found it in two more, flagged by no solver: a supplied format description reproducing four rows of its own reference, and worked examples that were themselves graded answers. During one frozen run the operator became the defect, misreading a slow transport as a broken one and instructing two running solvers to truncate; the one that complied lost six of thirteen questions, and the one that refused reported the misdiagnosis. A freeze constrains artefacts and says nothing about the operator's conduct during a run.

Version 1.1.0 adds a defect that was reachable by neither detector. A graded question scored a field the task statement never mentioned, so no solver could report its omission, and the reference was not itself wrong, so no check flagged it. Our adversarial suite had already constructed that exact submission and classified it as an attack, scoring it 94.6/100 — the same score, on the same question, that all six honest solvers received. This sharpens the paper's central claim: the three checks are functions of the reference, but the task statement is not, so none of them can detect a disagreement between the two. A benchmark can be internally consistent to the last decimal place and still be grading answers to a question it did not ask.