Benchmark and paper package for validating operational self-model and agency constructs in controlled artificial neural architectures.
-
Updated
Jul 6, 2026 - Python
Benchmark and paper package for validating operational self-model and agency constructs in controlled artificial neural architectures.
Construct-validity audit of the standard blood–brain barrier (BBB) peptide benchmark: an identity-controlled re-evaluation + shared-source provenance/overlap map, with an open, CPU-reproducible evaluation harness. Do these predictors measure penetration, or their benchmarks?
Inspect eval: do LLMs writing case notes separate observation from interpretation, and can they evade a lexical validator? Deterministic grader, pre-registered rubric.
Ecological study on administrative diabetes indicators and the care cascade using NDB Open Data, Japan (335 secondary medical areas, FY2023-2024)
Construct-validity diagnostics for mechanistic interpretability probes: demonstrating why held-out probe accuracy cannot establish semantic representation under confounded stimulus designs.
Add a description, image, and links to the construct-validity topic page so that developers can more easily learn about it.
To associate your repository with the construct-validity topic, visit your repo's landing page and select "manage topics."