Skip to content

huggingface/inspect_evals

Collection of evals for Inspect AI

https://skillcdn.ai/gh/huggingface/inspect_evals

Connect this address to your AI to use this skill. How to connect

  • Unverified
  • Default branchmain
  • Commitc876ff7
  • LicenseMIT
View on GitHub

ci-maintenance-workflow

CI and GitHub Actions maintenance workflows — fix a failing test from a CI URL, add @pytest.mark.slow markers to slow tests, or review a PR against agent-checkable standards. Use when user asks to fix a failing test, mark slow tests, or review a PR. Trigger when the user asks you to run the "Write a PR For A Failing Test", "Mark Slow Tests", or "Review PR According to Agent-Checkable Standards" workflow.

Path
.claude/skills/ci-maintenance-workflow/SKILL.md
License
MIT

CI Maintenance Workflows

This skill covers three CI/GitHub Actions-related workflows:

  • Fix A Failing Test: Diagnose and fix a failing test from a CI run
  • Mark Slow Tests: Add @pytest.mark.slow markers to tests exceeding thresholds
  • Review PR: Automatically check a PR against agent-checkable standards

Write a PR For A Failing Test

  1. The user should provide you with a link that looks like https://github.com/ArcadiaImpact/inspect-evals-actions/actions/runs/<number>/job/<number2>, or a test to fix. If they haven't, please ask them for the URL of the failing Github action, or a test that should be fixed.
  2. If a link is provided, you will be unable to view this link directly, but you can use 'gh run view' to see it.
  3. Checkout main and do a git pull. Make sure you're starting from a clean slate and not pulling in any changes from whatever you or the user was previously doing.
  4. Create a branch agent/<number> for your changes, in inspect_evals. Note this is not the same as inspect-evals-actions where the above URL was from.
  5. Identify the failing test(s) and any other problems.
  6. Attempt these fixes one by one and verify they work locally, with uv run pytest tests/<eval> --runslow.
  7. If your fix changes datasets, expected values, scoring logic, or anything that affects evaluation results, read TASK_VERSIONING.md and PACKAGE_VERSIONING.md, and make all appropriate changes.
  8. Perform linting as seen in .github/workflows/build.yml, and fix any problems that arise.
  9. Commit your changes to the branch.
  10. Before pushing, go over your changes and make sure all the changes that you have are high quality and that they are relevant to your finished solution. Remove any code that was used only for diagnosing the problem.
  11. Push your changes.
  12. Use the Github API to create a pull request with the change. Include a link to the failing test, which test failed, and how you fixed it.

TODO in future: Automate kicking off this workflow. Allow the agent to check the status of its PR and fix any failing checks through further commits autonomously.


Mark Slow Tests

This workflow adds @pytest.mark.slow(seconds) markers to tests that exceed the 10-second threshold in CI but are not already marked. This is necessary because our CI pipeline checks for unmarked slow tests and fails the build if any test takes >= 15 seconds without a marker.

Workflow Steps
  1. Get the GitHub Actions URL: The user should provide a URL like https://github.com/ArcadiaImpact/inspect-evals-actions/actions/runs/<run_id> or https://github.com/ArcadiaImpact/inspect-evals-actions/actions/runs/<run_id>/job/<job_id>. If they haven't, please ask them for the URL of the failing Github action.

  2. Extract slow test information: Use the fetch_slow_tests.py script to get all slow tests and their recommended markers:

    uv run python agent_artefacts/scripts/fetch_slow_tests.py <url_or_run_id> --threshold 10
    

    This script:

    • Fetches logs from all jobs in the GitHub Actions run
    • Parses test timing information from the "Check for unmarked slow tests" step
    • Aggregates durations across all Python versions and platforms
    • Reports the maximum duration observed for each test
    • Provides recommended @pytest.mark.slow(N) marker values (ceiling of max duration)

    For JSON output (useful for programmatic processing):

    uv run python agent_artefacts/scripts/fetch_slow_tests.py <url_or_run_id> --threshold 10 --json
    

    Tests marked [FAIL] in the output exceeded 60 seconds and caused the CI to fail. All tests above the threshold (default 10s) should be marked.

  3. Check existing markers: For each identified test, check if it already has a @pytest.mark.slow marker:

    grep -E "def <test_name>|@pytest\.mark\.slow" tests/<path>
    

    If the test already has a slow marker with a value >= the observed duration, skip it (per user instruction: "If the number is there but too high, leave it as is").

  4. Add slow markers: For tests without markers, add @pytest.mark.slow(N) where N is the observed maximum duration rounded up to the nearest integer. Place the marker immediately before other decorators like @pytest.mark.huggingface or @pytest.mark.asyncio.

    Example:

    # Before
    @pytest.mark.huggingface
    def test_example():
    
    # After
    @pytest.mark.slow(45)
    @pytest.mark.huggingface
    def test_example():
    
  5. Run linting: After adding markers, run linting to ensure proper formatting:

    uv run ruff check --fix <modified_files>
    uv run ruff format <modified_files>
    
  6. Commit: Commit the changes with a descriptive message listing all tests that were marked:

    Add slow markers to tests exceeding 10s threshold
    
    Tests marked:
    - <test_path>::<test_name>: slow(N)
    - ...
    
    Co-Authored-By: Claude <assistant_name> <noreply@anthropic.com>
    
Notes
  • The slow marker value should be slightly higher than the observed maximum to account for CI variance. Rounding up to the nearest integer is usually sufficient.
  • Parametrized tests (e.g., test_foo[param1]) share a single marker on the base test function. You should use the highest value found across all params for the slow marker.

TODO: What if one param is an order of magnitude higher than the rest? Should we recommend splitting it into its own test? pytest.mark.slow contents aren't that crucial, but still.


Review PR According to Agent-Checkable Standards

This workflow is designed to run in CI (GitHub Actions) to automatically review pull requests against the agent-checkable standards in EVALUATION_CHECKLIST.md. It produces a SUMMARY.md file that is posted as a PR comment.

Important Notes
  • This workflow is optimized for automated CI runs, not interactive use.
  • The SUMMARY.md file location is /tmp/SUMMARY.md (appropriate for GitHub runners).
  • Unlike other workflows, this one does NOT create folders in agent_artefacts since it runs in ephemeral CI environments.
Workflow Steps
  1. Identify if an evaluation was modified: Look at the changed files in the PR to determine which evaluation(s) are affected. Use git diff commands to see what has changed. If an evaluation (found in src/inspect_evals) was modified, complete all steps in the workflow. If no evaluation was modified, skip any steps that say (Evaluation) in them.

  2. Read the repo context: Read agent_artefacts/repo_context/REPO_CONTEXT.md if present. Use these lessons in your analysis.

  3. (Evaluation) Read the standards: Read EVALUATION_CHECKLIST.md, focusing on the Agent Runnable Checks section. These are the standards you will check against.

  4. (Evaluation) Check each standard: For each evaluation modified in the PR, go through every item in the Agent Runnable Checks section:

    • Read the relevant files in the evaluation
    • Compare against the standard
    • For any violations found, add an entry to /tmp/NOTES.md with:
      • Standard: Which standard was violated
      • Issue: Description of the issue
      • Location: file_path:line_number or file_path if not line-specific
      • Recommendation: What should be changed

    Only worry about violations that exist within the code that was actually changed. Pre-existing violations that were not touched or worsened by this PR can be safely ignored.

  5. Examine test coverage: Use the ensure-test-coverage skill and appropriate commands to identify any meaningful gaps in test coverage. You can skip this step if a previous step showed no tests were found, since this issue will already be pointed out.

  6. Apply code quality analysis: Examine all files changed for quality according to BEST_PRACTICES.md. Perform similar analysis and note-taking as in Step 5. If you notice poor quality in a way that is not mentioned in BEST_PRACTICES.md, add this to your notes, and under Standard, write "No explicit standard - agent's opinion". Err on the side of including issues here, unless REPO_CONTEXT.md or another document tells you not to - we will be updating these documents over time, and it is easier for us to notice something too nitpicky than to notice the absence of something important.

  7. Write the summary: After checking all standards, create /tmp/SUMMARY.md by consolidating the issues from NOTES.md:

    ## Summary
    
    Brief overview of what was reviewed and overall assessment.
    
    ## Issues Found
    
    ### [Standard Category Name]
    
    **Issue**: Description of the issue
    **Location**: `file_path:line_number` or `file_path` if not line-specific
    **Recommendation**: What should be changed
    
    (Repeat for each issue from NOTES.md)
    
    ## Notes
    
    Any additional observations or suggestions that aren't violations but could improve the code.
    

    If no issues are found, write a brief summary confirming the PR passes all agent-checkable standards. Do NOT include a list of passed checks - it is assumed that anything not mentioned is fine.

  8. Important: Always write to /tmp/SUMMARY.md even if there are no issues. The CI workflow depends on this file existing.

End your comment with 'This is an automatic review performed by Claude Code. Any issues raised here should be fixed or justified, but a human review is still required in order for the PR to be merged.'

What This Workflow Does NOT Check
  • Human-required checks (these need manual review)
  • Evaluation report quality (requires running the evaluation)
  • Whether the evaluation actually works (requires execution)

This workflow only checks code-level standards that can be verified by reading the code.

Files of this skill

This skill has no supporting files.

Browse all supporting files