Skip to content

fix(eval): parse code fences like lib.sh in the outputs eval - #1055

Open
Tong-bit-art wants to merge 1 commit into
msitarzewski:mainfrom
Tong-bit-art:fix/eval-fence-model
Open

Tong-bit-art wants to merge 1 commit into
msitarzewski:mainfrom
Tong-bit-art:fix/eval-fence-model

Conversation

@Tong-bit-art

Copy link
Copy Markdown
Contributor

What does this PR do?

test-convert-outputs.sh's split-integrity model (fence_blocks) read code
fences only at column 0, while CommonMark — and lib.sh's
fence_open_p / fence_closes_p, aligned to it in #855 and #1028 — allows up
to three spaces of indentation. An indented opening fence was therefore
invisible to the eval, its column-0 closer became the model's opener, and
the following lines were reported as a torn fenced block across
SOUL.md/AGENTS.md
:

  FAIL openclaw: fence-model-fixture source fenced block at lines 4-6 is torn
       across SOUL.md/AGENTS.md — a `## ` heading inside the fence was treated
       as a section boundary
  Results: 30 passed, 1 failed
  FAILED

The source is valid: GitHub renders the block, lint-agents.sh accepts it,
and convert_openclaw() splits it correctly (the model's "block at lines 4-6"
is its own mis-pairing, not the real three-line fence). Because the Check
Tools workflow runs this eval as a hard gate on every PR, a contributor
adding such an agent gets a failure that says their fence is torn when it is
not.

The model now mirrors lib.sh exactly: up to three spaces of indentation,
same character, a closer at least as long as the opener with nothing but
whitespace after the run, and an info-string run inside a block stays content.

Evidence

Reproduced in a throwaway repo with one fixture:

---
name: Fence Model Fixture
description: Fixture agent with an indented opening code fence
color: blue
---
## Identity
  ```text
  code sample
```
## Core Mission
mission text

On current main the eval fails with the report above; with this change all 15
tools validate. The existing roster is unchanged: the full eval reports
32 passed, 0 failed (282 agents × 15 tools) with zero manifest drift.

Tests

New scripts/test-convert-fence-model.sh builds the throwaway repo around the
real eval and that fixture, runs it with --update, and requires a clean
validation with no torn-block report. It fails on the previous model with the
false positive above and passes here.

Validation on the exact commit:

  • full bash suite 19/19 and Hermes plugin checks 3/3;
  • test-lint-fences.sh, strict test-convert-outputs.sh (zero drift),
    bash -n, git diff --check: clean.

Scope notes

  • Two existing files plus one new test; no generated output, no manifest
    change, no converter behavior change.
  • The eval's model is the only thing that changes: lib.sh and
    convert_openclaw() already handled indented fences.
  • One step is added to the existing Check Tools workflow to run the new
    regression; no new workflow.

AI assistance was used in preparing this patch; the reproduction and test
runs above were executed locally.

test-convert-outputs.sh's Python fence model read fences only at column 0. With an indented opening fence and a column-0 closer - valid CommonMark that GitHub renders, lint accepts and convert_openclaw splits correctly - the model missed the opener and treated the closer as its own opener, so it reported a torn fenced block across SOUL.md/AGENTS.md and failed the Check Tools workflow on an agent file that was fine.

Mirror lib.sh's fence_open_p / fence_closes_p (up to three spaces, same character, a closer at least as long as the opener with nothing but whitespace after the run) so the eval sees the same blocks the converter does.

test-convert-fence-model.sh builds a throwaway repo with one such fixture and runs the real eval; on the previous model it fails with the false torn-block report, and with this change all tools validate. The full roster still passes with zero manifest drift.
TheRealVitja pushed a commit to TheRealVitja/agency-agents that referenced this pull request Oct 6, 2026
…nts)

Curated sync of the open upstream pull requests as of 2026-10-06:

- Script and CI fixes: msitarzewski#1055, msitarzewski#1030, msitarzewski#865, msitarzewski#860, msitarzewski#967, msitarzewski#870, msitarzewski#869, msitarzewski#889,
  msitarzewski#1056, msitarzewski#755, msitarzewski#868, msitarzewski#771, msitarzewski#867; ported msitarzewski#523, msitarzewski#512 and the permissions
  part of msitarzewski#790.
- Existing-agent fixes: msitarzewski#1033-msitarzewski#1052, msitarzewski#1053, msitarzewski#1054, msitarzewski#1023, msitarzewski#756-msitarzewski#759, msitarzewski#799,
  msitarzewski#805, msitarzewski#715, msitarzewski#752, msitarzewski#784, msitarzewski#793, msitarzewski#858, msitarzewski#812, msitarzewski#789, msitarzewski#1007.
- New agents: msitarzewski#702, msitarzewski#707, msitarzewski#731, msitarzewski#732, msitarzewski#764, msitarzewski#848, msitarzewski#859, msitarzewski#862, msitarzewski#863, msitarzewski#886,
  msitarzewski#908, msitarzewski#982-msitarzewski#985, msitarzewski#1031, msitarzewski#1032.
- Docs: msitarzewski#577, msitarzewski#743, msitarzewski#762, msitarzewski#785, msitarzewski#786, msitarzewski#815, msitarzewski#816.

Fixes found while integrating: Bash 3.2 guard for the msitarzewski#755 worker argv,
Windsurf re-conversion over a stale .windsurfrules, locale-independent
check-divisions.sh, agency- prefix handling in the outputs eval, and the
India Business Navigator's YAML and headings. Resolves upstream issues
msitarzewski#229, msitarzewski#763, msitarzewski#821 and msitarzewski#1027.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NwgfpJ9tGbUh5g84u5VSgv
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant