Official repository for DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
-
Updated
Jun 17, 2026 - Python
Official repository for DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Official implementation for the OpenRubric family: a line of work on rubric-based reward modeling for LLMs.
Human-reviewed workflow prototypes for rubric evidence, feedback drafting, reviewer packets, and remediation actions.
Prompt and AI-output evaluation case studies for hallucination risk, localization, rubrics, and QA.
AI-agent evaluation harness with rubrics, regression datasets, deterministic checks, and reports.
Rubric-as-judge runner using the Anthropic Python SDK with tool-use API for schema-validated structured output. 17 mocked tests included.
Local-first AI evaluation and robotics review tools: EvalKit, Robotics ReviewKit, release evidence, and technical buying resources.
Auto-generate evaluation rubrics from agent audit-log trajectories (PhoneWorld pattern applied to action logs)
Lightweight examples of LLM evaluation artifacts: rubrics, prompts, and evaluator guidelines.
Teacher-led, AI-assisted authoring of high-quality assessment questions from real teaching materials — with QTI export for Inspera and other QTI-compatible platforms.
Add a description, image, and links to the rubrics topic page so that developers can more easily learn about it.
To associate your repository with the rubrics topic, visit your repo's landing page and select "manage topics."