Repository navigation
fix(eval): parse code fences like lib.sh in the outputs eval - #1055
Open
Tong-bit-art wants to merge 1 commit into
Open
Tong-bit-art wants to merge 1 commit into
Tong-bit-art wants to merge 1 commit into
Conversation
test-convert-outputs.sh's Python fence model read fences only at column 0. With an indented opening fence and a column-0 closer - valid CommonMark that GitHub renders, lint accepts and convert_openclaw splits correctly - the model missed the opener and treated the closer as its own opener, so it reported a torn fenced block across SOUL.md/AGENTS.md and failed the Check Tools workflow on an agent file that was fine. Mirror lib.sh's fence_open_p / fence_closes_p (up to three spaces, same character, a closer at least as long as the opener with nothing but whitespace after the run) so the eval sees the same blocks the converter does. test-convert-fence-model.sh builds a throwaway repo with one such fixture and runs the real eval; on the previous model it fails with the false torn-block report, and with this change all tools validate. The full roster still passes with zero manifest drift.
TheRealVitja
pushed a commit
to TheRealVitja/agency-agents
that referenced
this pull request
Oct 6, 2026
…nts) Curated sync of the open upstream pull requests as of 2026-10-06: - Script and CI fixes: msitarzewski#1055, msitarzewski#1030, msitarzewski#865, msitarzewski#860, msitarzewski#967, msitarzewski#870, msitarzewski#869, msitarzewski#889, msitarzewski#1056, msitarzewski#755, msitarzewski#868, msitarzewski#771, msitarzewski#867; ported msitarzewski#523, msitarzewski#512 and the permissions part of msitarzewski#790. - Existing-agent fixes: msitarzewski#1033-msitarzewski#1052, msitarzewski#1053, msitarzewski#1054, msitarzewski#1023, msitarzewski#756-msitarzewski#759, msitarzewski#799, msitarzewski#805, msitarzewski#715, msitarzewski#752, msitarzewski#784, msitarzewski#793, msitarzewski#858, msitarzewski#812, msitarzewski#789, msitarzewski#1007. - New agents: msitarzewski#702, msitarzewski#707, msitarzewski#731, msitarzewski#732, msitarzewski#764, msitarzewski#848, msitarzewski#859, msitarzewski#862, msitarzewski#863, msitarzewski#886, msitarzewski#908, msitarzewski#982-msitarzewski#985, msitarzewski#1031, msitarzewski#1032. - Docs: msitarzewski#577, msitarzewski#743, msitarzewski#762, msitarzewski#785, msitarzewski#786, msitarzewski#815, msitarzewski#816. Fixes found while integrating: Bash 3.2 guard for the msitarzewski#755 worker argv, Windsurf re-conversion over a stale .windsurfrules, locale-independent check-divisions.sh, agency- prefix handling in the outputs eval, and the India Business Navigator's YAML and headings. Resolves upstream issues msitarzewski#229, msitarzewski#763, msitarzewski#821 and msitarzewski#1027. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NwgfpJ9tGbUh5g84u5VSgv
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
test-convert-outputs.sh's split-integrity model (fence_blocks) read codefences only at column 0, while CommonMark — and
lib.sh'sfence_open_p/fence_closes_p, aligned to it in #855 and #1028 — allows upto three spaces of indentation. An indented opening fence was therefore
invisible to the eval, its column-0 closer became the model's opener, and
the following lines were reported as a torn fenced block across
SOUL.md/AGENTS.md:
The source is valid: GitHub renders the block,
lint-agents.shaccepts it,and
convert_openclaw()splits it correctly (the model's "block at lines 4-6"is its own mis-pairing, not the real three-line fence). Because the Check
Tools workflow runs this eval as a hard gate on every PR, a contributor
adding such an agent gets a failure that says their fence is torn when it is
not.
The model now mirrors
lib.shexactly: up to three spaces of indentation,same character, a closer at least as long as the opener with nothing but
whitespace after the run, and an info-string run inside a block stays content.
Evidence
Reproduced in a throwaway repo with one fixture:
On current main the eval fails with the report above; with this change all 15
tools validate. The existing roster is unchanged: the full eval reports
32 passed, 0 failed (282 agents × 15 tools) with zero manifest drift.
Tests
New
scripts/test-convert-fence-model.shbuilds the throwaway repo around thereal eval and that fixture, runs it with
--update, and requires a cleanvalidation with no torn-block report. It fails on the previous model with the
false positive above and passes here.
Validation on the exact commit:
test-lint-fences.sh, stricttest-convert-outputs.sh(zero drift),bash -n,git diff --check: clean.Scope notes
change, no converter behavior change.
lib.shandconvert_openclaw()already handled indented fences.regression; no new workflow.
AI assistance was used in preparing this patch; the reproduction and test
runs above were executed locally.