benchmark-lightflows
Run and interpret the unified A/B Lightflow benchmark (benchmarks/run_benchmark.py) comparing Arm A (Lightflow DAG) against Arm B (Traditional Skill) across 7 matched domain workflow pairs (3 to 22 stages).
작성자를 위한 경고
- "name" should match the name of the directory that holds SKILL.md
- Not listed through the skills extension: the directory must be named after the skill.
Benchmarking Lightflow Workflows
Use benchmarks/run_benchmark.py to measure and verify Arm A (Lightflow DAG) against Arm B (Steelmanned Traditional Skill:
benchmarks/baseline_skills/<name>/SKILL.md + examples/<name>/actions.py)
across all 7 domain workflow pairs (hn_digest, usgs_seismic_alert,
pypi_upgrade_guard, async_job_watcher, blue_green_release,
incident_db_failover, tenant_gitops_onboarding) plus the create_lightflow
authoring workflow.
Commands
# Run the unified A/B benchmark (Markdown report)
python3 benchmarks/run_benchmark.py
# Run N=5 independent trials to verify multi-run consistency
python3 benchmarks/run_benchmark.py --trials=5
# Emit machine-readable JSON metrics
python3 benchmarks/run_benchmark.py --trials=5 --json
What Is Verified on Every Run
- Ground-Truth Correctness & Invariants:
order_match: Completed stages inpassport.jsonmatch the exact expected topological sequence.no_duplicate_side_effects: Every stage completes at most once (1x), including across mid-run failures andlightflow resumerecoveries (pypi_upgrade_guard,blue_green_release,incident_db_failover).gate_discipline: Everyoperator_actionstage records aPAUSEDstamp (exit 2) before itsCOMPLETEDstamp.external_state_valid&baseline_external_state_valid: Both Arm A and Arm B produce the exact expected on-disk artifacts, intermediate rollbacks, and cleanups.
- Context & Code Generation Cost (
1 token ≈ 4 chars):- Skill Loaded: Universal
skills/run_lightflows/SKILL.md(~1,097 tokonce per session) vs. per-workflowbenchmarks/baseline_skills/<name>/SKILL.md(~7,075 tokacross the 7 domain workflows). - Agent Command Output (
cmd): Compact declarativelightflow start / resumeCLI calls (~795 tokacross 7 workflows) vs. inlinepython3 -corchestration snippets (~2,992 tokacross 7 workflows). - Out-of-Context State:
passport.jsonbytes persisted on disk (52,958 chars/~13,240 tok) andlightflow.yaml + actions.pysource bytes avoided (115,321 chars/~28,830 tok, 83% warm zero-context savings).
- Skill Loaded: Universal