0
feat(signals): report research steering eval suite
Source: PostHog/posthog#91313 · opened by @andrewm4894
What Build an eval suite for the report research stage's steering behavior — the Steering from this team section that #91311 (closing #91267) renders into the research prompt. Nothing verifies that behavior today. products/signals/eval/ holds only eval_grouping_e2e.py and eval_scout_safety.py; neither exercises report research. This is a suite to write from scratch, not a case to add — eval_scout_safety.py (the one existing eval that tests an agent resisting hostile input) and conftest.py are the patterns to copy. Why #91267's acceptance criteria required the steering loop to be eval-proven, and #91311 deliberately shipped without it, deferring the suite to its own PR — this issue is that PR's tracking. The stakes rose with #91311's design: the research prompt now carries every note origin, including the derived ones (report_dismissal, report_discussion, report_feedback) whose text is built from raw product data. The only guard against a hostile note steering …
No pledges yet. Be the first to back this.
Comments
Similar requests
feat(signals): report research agent reads steering notes and fleet memory
0 votes · 0 comments
feat(signals): pipeline audience targets for steering notes
0 votes · 0 comments
feat(signals): report research agent remembers judgments in the fleet scratchpad
0 votes · 0 comments
feat(tasks): implementation runs read steering notes and fleet memory
0 votes · 0 comments
feat(signals): comment on the source GitHub issue when an inbox report starts tracking it
0 votes · 0 comments
No comments yet.