This is an automated email from the ASF dual-hosted git repository.

davsclaus pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/camel-performance-tests.git


The following commit(s) were added to refs/heads/main by this push:
     new fcc65cc  The stepwise ladder: HTTP rungs, tool groups and a 
final-state score
fcc65cc is described below

commit fcc65ccfdf84d81dcffeb0188ca92900a9e8db71
Author: Claus Ibsen <[email protected]>
AuthorDate: Mon Oct 5 16:07:14 2026 +0200

    The stepwise ladder: HTTP rungs, tool groups and a final-state score
    
    - run-http.sh runs the four HTTP rungs; ref_pass.py and ref-server.sh run 
the reference pass and the reference stock
      API server.
    - agent_mcp_stepwise.py: a 64k context by default (BENCH_NUM_CTX); 
BENCH_TOOL_GROUPS=1 lets the model ask for the
      tool groups the running app offers (CAMEL-24834); the camel_runtime_* 
tools get the app's name when the model
      leaves it out.
    - summarize_stepwise.py reports "final state right" next to the strict 
score, as a model that tests its own app
      with HTTP requests logs errors on the way to a right answer.
    - gen_stepwise.py: the stock lookups are written in Groovy instead of 
JSONPath; the openapi-server, json-transform
      and order-lines requests are worded so a careful reader cannot misread 
them.
    - The README describes the stepwise ladder: where the steps come from, what 
a step is scored on, the tools and
      model settings, the output and the pitfalls.
    
    Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
    Claude-Session: https://claude.ai/code/session_01STT6whBgK1AqsSsUKrnE8m
---
 ai-benchmark/.gitignore                            |  2 +
 ai-benchmark/README.md                             | 97 +++++++++++++++++++++-
 ai-benchmark/agent_mcp_stepwise.py                 | 26 +++++-
 ai-benchmark/gen_stepwise.py                       | 36 ++++----
 ai-benchmark/ref-server.sh                         | 17 ++++
 ai-benchmark/ref_pass.py                           | 36 ++++++++
 ai-benchmark/run-http.sh                           |  8 ++
 ai-benchmark/run-stepwise-ladder.sh                |  4 +-
 .../_peer-openapi-server/openapi-server.camel.yaml | 25 ++----
 ai-benchmark/summarize_stepwise.py                 | 17 ++--
 10 files changed, 218 insertions(+), 50 deletions(-)

diff --git a/ai-benchmark/.gitignore b/ai-benchmark/.gitignore
index 4a84651..183f231 100644
--- a/ai-benchmark/.gitignore
+++ b/ai-benchmark/.gitignore
@@ -1,6 +1,7 @@
 # run outputs
 oneshot-*/
 stepwise/
+ref/
 local/
 *.log
 *.out
@@ -11,4 +12,5 @@ stepwise-ladder/
 stepwise-project/
 # per-run notes and summaries written beside a run
 *.build.txt
+*.jars.sha
 *.summary
diff --git a/ai-benchmark/README.md b/ai-benchmark/README.md
index b65287c..3255cee 100644
--- a/ai-benchmark/README.md
+++ b/ai-benchmark/README.md
@@ -22,7 +22,8 @@ Two benchmarks, both scored by what actually runs, not by 
reading the model's ou
 2. **Stepwise** (`agent_mcp_stepwise.py`): the model edits a running 
integration one request at a time ("fire every
    five seconds", "add a choice", "add error handling", eight requests) 
through the write, run, log and error tools of
    the MCP server. Each step is scored on the file on disk (a regex), the 
properties file, the log of the running
-   integration after the reload, and the errors, against a reference 
checkpoint.
+   integration after the reload, and the errors, against a reference 
checkpoint. Since round 2 the stepwise benchmark
+   runs over a ladder of examples, one steps file each; see [The stepwise 
ladder](#the-stepwise-ladder).
 
 The model is never shown the example projects or their catalog tools; the 
one-shot tool set is restricted to catalog
 lookups and validation (`BENCH_TOOL_ALLOW` in `run-suite.sh`): the shared 
`camel_catalog_doc`, `camel_catalog_find`,
@@ -127,3 +128,97 @@ Added 2026-09-17 for the second series.
   - `pre` / `post`: shell commands run in this directory around the example 
(openapi-client starts the reference
     petstore server from `seed/openapi-server-ref/` and stops it after).
 - Docker Desktop must be running for `infra`; `camel infra` pulls the images 
on first use.
+
+## The stepwise ladder
+
+Added in round 2 (2026-09-19 onwards); the main benchmark since 2026-09-21, 
because building an integration one
+request at a time is how people work with Camel, and the one-shot prompt is 
not.
+
+Each rung of the 
[camel-jbang-examples](https://github.com/apache/camel-jbang-examples) ladder 
that runs without
+Docker (plus the sql rung, with `camel infra`) is one stepwise scenario, 
written from its README's "Build it step by
+step" section: the README's step 1 is the starting project, and each
+later step is one request to the model. The model edits the running 
integration through the MCP server; the harness
+scores each step before sending the next.
+
+```bash
+./start-server.sh                       # set CAMEL_MCP_JAR to measure a 
locally built camel-jbang-mcp
+python3 gen_stepwise.py                 # writes steps-ladder/<name>.json and 
stepwise-ladder/<name>/ for every example
+BENCH_REFERENCE=1 ./run-stepwise-ladder.sh ref       # the reference pass: 
must pass every step before a model run
+./run-stepwise-ladder.sh s1 5           # five runs: stepwise/s1-1 .. s1-5, 
then the summary table
+./run-http.sh h1 3                      # only the four HTTP rungs: stock-api, 
http-client, openapi-server/-client
+BENCH_ONLY="transform-xslt route-aggregator" ./run-stepwise-ladder.sh x 1   # 
any subset, names as in steps-ladder/
+python3 summarize_stepwise.py s1 5      # the summary again: passes per 
example and step, runs with every step passed
+```
+
+A full ladder run of the 21 examples takes about 65 minutes on the Mac mini 
the series runs on (qwen3.6:35b-a3b in
+Ollama); the four HTTP rungs about 13. Every JSON
+file in `steps-ladder/` is run, so keep hand-made scenario files out of it.
+
+### Where the steps come from
+
+`gen_stepwise.py` holds the definitions: for each example the starting files 
(`initial`), the steps (`request`,
+`check`, `reference`) and the settings below. `steps-ladder/` and 
`stepwise-ladder/` are generated (git ignores them),
+so change a step in `gen_stepwise.py` and never in the JSON. The runner calls 
`gen_stepwise.py <name>` before every
+pass, which rewrites that example's steps file and resets its project, so each 
pass starts from the same files.
+
+- `seed`: files under `seed/<name>/` copied into the project (data files such 
as `stock.json`, `orders/`, a contract);
+  `exclude_seeds` leaves out a file the model is asked to write 
(`packing-slip.xsl`).
+- `peer`: a second app the example talks to, started with `camel run 
--source-dir` before the model's app and stopped
+  after it. `contracts-openapi-client` calls the stock API in 
`seed/_peer-openapi-server` on port 8080.
+- `infra`: services started with `camel infra run` (the sql rung's Postgres); 
they stay up across the passes of a run
+  and are stopped when the run is over. Docker must be running.
+- `restart`: on a step that adds or changes Java, the harness restarts the app 
after the model's turn, as the README's
+  step says; a reload does not compile a class.
+- `wait_seconds`: how long to wait after the reload before the log and files 
are checked.
+
+### What a step is scored on
+
+`ok` means `file_ok`, `props_ok` and `log_ok` all hold and the step logged no 
new errors (`ok_final` is the same
+without the errors condition; the summary reports it as "final state right", 
since a model that tests its own app
+with the HTTP request tool causes errors on the way to a right answer, which 
`ok` counts against it). The checks, all
+optional:
+
+| Check | Passes when |
+|---|---|
+| `file_regex`, `file_regex2` | the route file matches (both, if both are 
given) |
+| `file_not_regex` | the route file does not match |
+| `props_regex` | `application.properties` matches |
+| `files` | `{name: regex}`: another project file matches; found by name 
anywhere in the project if not at that path |
+| `log_regex` | a regex or a list of them; every one matches an INFO or WARN 
record (`min_log` times, default 1) |
+| `log_not_regex` | no record logged after this step's reload matches |
+| `interval_min` | the first two records matching `log_regex` are at least 
this many seconds apart |
+| `probes` | HTTP requests (`method`, `url`, `headers`, `body`, 
`expect_status`, `body_regex`) sent after the reload; each is retried up to six 
times, two seconds apart, until status and body match, since the old route may 
still answer during a reload |
+| `errors_ok` | errors in the log do not fail the step (a step that provokes 
failures on purpose) |
+| `min_errors` | at least this many errors must be logged |
+
+Errors are counted per step: ERROR records in the log and entries from 
`camel_get_errors` that were not there before
+the step.
+
+### Tools and model settings
+
+The model gets the shared authoring tools (`SHARED` in 
`agent_mcp_stepwise.py`: catalog doc/find/sample, validate,
+get/write/edit file, run, control, log, errors, eval expression, error 
diagnose) plus `camel_catalog_docs`,
+`camel_component_properties` and `camel_configuration_validate`, and at most 
20 tool calls per step
+(`BENCH_TOOL_CALLS`), in a 64k context (`BENCH_NUM_CTX`; at 32k the longer 
examples ran out of context by their last
+step and the model stopped without a tool call). `camel_edit_file` is in the 
set because a model that rewrites a whole
+file to change one line corrupts lines it was not asked to touch. 
`BENCH_EXTRA_TOOLS` adds tools for an experiment
+(the SQL tool group of CAMEL-24834).
+
+The harness sends `think: false` to Ollama. With thinking on, the HTTP series 
s5 lost 4 of 33 steps to thinking
+spirals of 4 to 8 minutes and 12k to 28k tokens that ended without an answer.
+
+### What comes out
+
+`stepwise/<tag>/<name>/`: `results.json` (per step: the checks, probe answers, 
errors with their first lines, tool
+calls, refused writes, seconds, tokens, the model's closing answer), `run.log` 
(one line per step), `step<n>.trace.jsonl`
+(every model turn and tool call), `step<n>.after.yaml` (the route file after 
the step), `camel-run.out` and
+`peer-run.out` (the apps' console). `stepwise-<tag>.log` has one line per 
example.
+
+### Pitfalls
+
+- Never write a log or output file inside a project the app watches 
(`--source-dir`): every write reloads the app,
+  and a reload that logs writes again. Output belongs under `stepwise/`, which 
is outside the projects.
+- Run the reference pass after any change to `gen_stepwise.py`, a seed, or 
Camel: a reference that fails is a
+  harness or Camel bug, not a model result.
+- Do not share one MCP server between two runs, a reference pass included: the 
step logs and error counts of one get
+  mixed into the other.
diff --git a/ai-benchmark/agent_mcp_stepwise.py 
b/ai-benchmark/agent_mcp_stepwise.py
index 14c72bc..c185c0d 100755
--- a/ai-benchmark/agent_mcp_stepwise.py
+++ b/ai-benchmark/agent_mcp_stepwise.py
@@ -24,11 +24,17 @@ HERE = os.path.dirname(os.path.abspath(__file__))
 TAG = os.environ.get("BENCH_TAG", "mcp-" + MODEL.replace(":", 
"_").replace("/", "_"))
 OUT = os.path.join(HERE, "stepwise", TAG)
 MAX_TOOL_CALLS = int(os.environ.get("BENCH_TOOL_CALLS", "20"))
+# the conversation grows over an example's steps (logs with stack traces are 
the bulk); at 32k a long example ran out
+# of context and the model stopped after a few words without a tool call 
(round 3, circuit-breaker steps 2-3)
+NUM_CTX = int(os.environ.get("BENCH_NUM_CTX", "65536"))
 # how many log records a step is scored on. A step that provokes failures on 
purpose (errors_ok) fills the window
 # with its own expected errors, and the line the check looks for scrolls out 
of it while sitting in the run log
 LOG_WINDOW = int(os.environ.get("BENCH_LOG_WINDOW", "400"))
 # round 2: BENCH_REFERENCE=1 skips the model and applies each step's reference 
files instead, to check the steps file itself
 REFERENCE = os.environ.get("BENCH_REFERENCE") == "1"
+# CAMEL-24834: before each step, ask camel_runtime_tool_groups what the app 
has and offer its groups' tools and guidance,
+# as the AI panel of the camel-jbang views does for a local model (the tool 
list changes only when the fingerprint does)
+TOOL_GROUPS = os.environ.get("BENCH_TOOL_GROUPS") == "1"
 TOOL_RESULT_CAP = 6000
 
 SHARED = ["camel_catalog_doc", "camel_catalog_find", "camel_catalog_sample", 
"camel_validate_source", "camel_get_files", "camel_write_file",
@@ -62,7 +68,7 @@ def ollama_chat(messages, tools):
     # thinking off, as the Camel CLI's own Ollama client sends and as the 
one-shot harness does: with it on, four of
     # 33 HTTP steps in series s5 spent 4-8 minutes and 12-28k tokens and 
answered nothing (CAMEL-24886)
     body = json.dumps({"model": MODEL, "messages": messages, "tools": tools, 
"stream": False, "think": False,
-                       "options": {"temperature": 0.2, "num_ctx": 
32768}}).encode()
+                       "options": {"temperature": 0.2, "num_ctx": 
NUM_CTX}}).encode()
     req = urllib.request.Request(HOST + "/api/chat", data=body, 
headers={"Content-Type": "application/json"})
     t0 = time.time()
     with urllib.request.urlopen(req, timeout=1800) as r:
@@ -396,11 +402,23 @@ def main():
         name = (data or {}).get("name") or os.path.basename(project)
         time.sleep(6)
 
-    messages = [{"role": "system", "content": SYSTEM.replace("{directory}", 
project).replace("{name}", name)}]
+    base_system = SYSTEM.replace("{directory}", project).replace("{name}", 
name)
+    messages = [{"role": "system", "content": base_system}]
+    groups_fp = None
     results = []
     try:
         for step in cfg["steps"]:
             sid = step["id"]
+            if TOOL_GROUPS and not REFERENCE:
+                data, raw = jcall(mcp, "camel_runtime_tool_groups", 
{"nameOrPid": name})
+                if isinstance(data, dict) and data.get("fingerprint") != 
groups_fp:
+                    groups_fp = data.get("fingerprint")
+                    extra = [t for g in data.get("groups") or [] for t in 
g.get("tools") or []]
+                    tools = to_ollama_tools(all_tools, SHARED + EXTRA + [t for 
t in extra if t not in SHARED + EXTRA])
+                    guidance = [g["guidance"] for g in data.get("groups") or 
[] if g.get("guidance")]
+                    messages[0]["content"] = base_system + (
+                        "\nThe running integration:\n" + "".join("- " + g + 
"\n" for g in guidance) if guidance else "")
+                print(f"step{sid}: tool groups {groups_fp} -> {raw[:300]}", 
file=log, flush=True)
             before = {"route": read(os.path.join(project, cfg["route_file"])), 
"props": read(os.path.join(project, cfg.get("props_file", 
"application.properties")))}
             before["error_records"], before["error_count"], 
before["reload_records"], before["log_records"] \
                 = error_snapshot(mcp, name)
@@ -435,6 +453,10 @@ def main():
                             args["directory"] = project
                         if fn["name"] in ("camel_get_log", "camel_get_errors", 
"camel_control", "camel_eval_expression") and "name" not in args:
                             args["name"] = name
+                        # the runtime tools take nameOrPid; the AI panel 
always means the selected integration, and
+                        # without it they fail on "multiple processes running" 
when a peer app runs beside it
+                        if fn["name"].startswith("camel_runtime_") and not 
args.get("nameOrPid"):
+                            args["nameOrPid"] = name
                         if fn["name"] in ("camel_write_file", 
"camel_edit_file"):
                             writes += 1
                         try:
diff --git a/ai-benchmark/gen_stepwise.py b/ai-benchmark/gen_stepwise.py
index 28213b7..9de521e 100644
--- a/ai-benchmark/gen_stepwise.py
+++ b/ai-benchmark/gen_stepwise.py
@@ -8,7 +8,7 @@ steps 2..N are the requests. Checks: log_regex (a regex or a 
list, all must matc
 file_regex (the route file), props_regex, files {name: regex}, errors_ok, 
min_errors.
 
 Usage: gen_stepwise.py            writes every steps file and project
-       gen_stepwise.py <name>     rewrites one project directory from its 
steps file (the runner does this per run)
+       gen_stepwise.py <name>     rewrites one example's steps file and resets 
its project (the runner does this per pass)
 """
 import datetime, json, os, shutil, sys
 
@@ -157,7 +157,7 @@ EXAMPLES.append({
     "seed": "transform-json-transform",
     "initial": {JT: route("json-transform", TIMER1, BODY_FILE + log("Order: 
${body}")), "application.properties": ""},
     "steps": [
-        {"request": "Read the order id into a header orderId with a jsonpath 
expression $.orderId and log \"Order <id>: <body>\".",
+        {"request": "Read the order id into a header orderId with a jsonpath 
expression $.orderId and log the message \"Order ${header.orderId}: ${body}\" 
(it reads Order ORD-1001: followed by the order JSON).",
          "check": {"file_regex": "jsonpath", "log_regex": "Order ORD-1001: 
\\{"},
          "reference": {JT: route("json-transform", TIMER1, BODY_FILE + H_ID + 
log("Order ${header.orderId}: ${body}"))}},
         {"request": "Add the number of lines as a header itemCount with 
jsonpath $.lines.length() and log \"Order <id> with <count> lines: <body>\".",
@@ -260,7 +260,7 @@ EXAMPLES.append({
     "seed": "route-order-lines",
     "initial": {OL: route("order-lines", FILE_ORDERS, UNMARSHAL + log("Order 
${body[orderId]} with ${body[lines].size()} line(s)")), 
"application.properties": ""},
     "steps": [
-        {"request": "Add a split over the order's lines (${body[lines]}) and 
inside it log \"  pick <qty> x <sku>\" for each line.",
+        {"request": "Add a split over the order's lines (${body[lines]}) and 
inside it log the message \"  pick ${body[qty]} x ${body[sku]}\" for each line 
(it reads   pick 2 x CAMEL-TSHIRT).",
          "check": {"file_regex": "split", "log_regex": "pick 2 x 
CAMEL-TSHIRT"},
          "reference": {OL: ol_s2}},
         {"request": "Keep the order id in a header orderId before the split 
and use it in the pick line: \"  pick <qty> x <sku> for <orderId>\".",
@@ -603,15 +603,14 @@ def set_header(name, kind, expr, ind="        "):
             + ind + "        expression: " + expr + "\n")
 CT_JSON = set_header("Content-Type", "constant", "application/json")
 DIRECT = lambda n: "      uri: direct\n      parameters:\n        name: " + n 
+ "\n"
-JSONPATH_SKU = ("        - setBody:\n            expression:\n              
jsonpath:\n                expression: \"$[?(@.sku == '${header.sku}')]\"\n"
-                "                resultType: java.util.List\n")
-NOT_FOUND = ("        - choice:\n            when:\n              - 
expression:\n                  simple:\n                    expression: 
\"${body.size()} == 0\"\n"
+GROOVY_SKU = ("        - unmarshal:\n            json:\n              library: 
Jackson\n"
+              "        - setBody:\n            expression:\n              
groovy:\n                expression: \"body.find { it.sku == headers.sku }\"\n")
+NOT_FOUND = ("        - choice:\n            when:\n              - 
expression:\n                  simple:\n                    expression: 
\"${body} == null\"\n"
              "                steps:\n" + set_header("CamelHttpResponseCode", 
"constant", "\"404\"", "                  ")
              + "                  - setBody:\n                      
expression:\n                        simple:\n                          
expression: '{\"error\": \"unknown sku ${header.sku}\"}'\n"
-             "            otherwise:\n              steps:\n                - 
setBody:\n                    expression:\n                      simple:\n      
                  expression: \"${body[0]}\"\n"
+             "            otherwise:\n              steps:\n"
              "                - marshal:\n                    json:\n          
            library: Jackson\n")
-FIRST_ITEM = ("        - setBody:\n            expression:\n              
simple:\n                expression: \"${body[0]}\"\n"
-              "        - marshal:\n            json:\n              library: 
Jackson\n")
+FIRST_ITEM = ("        - marshal:\n            json:\n              library: 
Jackson\n")
 
 # ---------------------------------------------------------------- 
connect/stock-api (a server: the checks are the README's curl)
 SA = "stock-api.camel.yaml"
@@ -620,8 +619,8 @@ REST_BOTH = "- rest:\n    path: /stock\n    get:\n      - 
to:\n          uri: di
 ALL_STOCK = route("all-stock", DIRECT("all-stock"), 
const("resource:file:stock.json") + CT_JSON)
 sa_s1 = REST_ALL + "\n" + ALL_STOCK
 sa_s2 = REST_BOTH + "\n" + ALL_STOCK + "\n" + route("one-sku", 
DIRECT("one-sku"), log("Stock asked for ${header.sku}") + 
const("resource:file:stock.json") + CT_JSON)
-sa_s3 = REST_BOTH + "\n" + ALL_STOCK + "\n" + route("one-sku", 
DIRECT("one-sku"), log("Stock asked for ${header.sku}") + 
const("resource:file:stock.json") + JSONPATH_SKU + FIRST_ITEM + CT_JSON)
-sa_s4 = REST_BOTH + "\n" + ALL_STOCK + "\n" + route("one-sku", 
DIRECT("one-sku"), log("Stock asked for ${header.sku}") + 
const("resource:file:stock.json") + JSONPATH_SKU + NOT_FOUND + CT_JSON)
+sa_s3 = REST_BOTH + "\n" + ALL_STOCK + "\n" + route("one-sku", 
DIRECT("one-sku"), log("Stock asked for ${header.sku}") + 
const("resource:file:stock.json") + GROOVY_SKU + FIRST_ITEM + CT_JSON)
+sa_s4 = REST_BOTH + "\n" + ALL_STOCK + "\n" + route("one-sku", 
DIRECT("one-sku"), log("Stock asked for ${header.sku}") + 
const("resource:file:stock.json") + GROOVY_SKU + NOT_FOUND + CT_JSON)
 GET = lambda path, status=200, rx=None: {"method": "GET", "url": 
"http://localhost:8080"; + path, "expect_status": status, **({"body_regex": rx} 
if rx else {})}
 EXAMPLES.append({
     "name": "connect-stock-api", "route_file": SA, "props_file": 
"application.properties", "wait_seconds": 4,
@@ -631,10 +630,10 @@ EXAMPLES.append({
         {"request": "Add a second operation to the rest: GET /stock/{sku}, 
routed to direct:one-sku, and a route one-sku that logs \"Stock asked for 
${header.sku}\" (the path parameter arrives as a header) and for now answers 
with the whole stock file like all-stock does.",
          "check": {"file_regex": "\\{sku\\}", "log_regex": "Stock asked for 
CAMEL-MUG", "probes": [GET("/stock/CAMEL-MUG", 200, "CAMEL-MUG")]},
          "reference": {SA: sa_s2}},
-        {"request": "In one-sku, pick the SKU out of the stock file with a 
jsonpath expression $[?(@.sku == '${header.sku}')] (resultType java.util.List), 
answer with the first element marshalled to JSON with Jackson, Content-Type 
application/json.",
-         "check": {"file_regex": "jsonpath", "probes": 
[GET("/stock/CAMEL-MUG", 200, "\"sku\":\"CAMEL-MUG\",\"qty\":42")]},
+        {"request": "In one-sku, keep loading the stock file into the body as 
now, then unmarshal the body with Jackson and pick the SKU with the Groovy 
expression body.find { it.sku == headers.sku } (the item as a map, or null), 
answer with the item marshalled to JSON with Jackson, Content-Type 
application/json.",
+         "check": {"file_regex": "groovy", "probes": [GET("/stock/CAMEL-MUG", 
200, "\"sku\":\"CAMEL-MUG\",\"qty\":42")]},
          "reference": {SA: sa_s3}},
-        {"request": "When the SKU is unknown (the jsonpath list is empty) 
answer with HTTP status 404 (the header CamelHttpResponseCode) and the body 
{\"error\": \"unknown sku ${header.sku}\"}; a known SKU still answers as 
before.",
+        {"request": "When the SKU is unknown (the Groovy expression gives 
null) answer with HTTP status 404 (the header CamelHttpResponseCode) and the 
body {\"error\": \"unknown sku ${header.sku}\"}; a known SKU still answers as 
before.",
          "check": {"file_regex": "404", "probes": [GET("/stock/CAMEL-SOCKS", 
404, "unknown sku CAMEL-SOCKS"), GET("/stock/CAMEL-MUG", 200, "\"qty\":42")]},
          "reference": {SA: sa_s4}},
     ]})
@@ -642,7 +641,7 @@ EXAMPLES.append({
 # ---------------------------------------------------------------- 
connect/http-client (the stock service is in the same app)
 HC = "http-client.camel.yaml"
 REST_ONE = "- rest:\n    path: /stock\n    get:\n      - path: \"/{sku}\"\n    
    to:\n          uri: direct:one-sku\n"
-STOCK_SERVICE = route("stock-service", DIRECT("one-sku"), 
const("resource:file:stock.json") + JSONPATH_SKU + NOT_FOUND + CT_JSON)
+STOCK_SERVICE = route("stock-service", DIRECT("one-sku"), 
const("resource:file:stock.json") + GROOVY_SKU + NOT_FOUND + CT_JSON)
 SERVER = REST_ONE + "\n" + STOCK_SERVICE + "\n"
 def set_prop(name, expr, ind="        "):
     return ind + "- setProperty:\n" + ind + "    name: " + name + "\n" + ind + 
"    expression:\n" + ind + "      simple:\n" + ind + "        expression: \"" 
+ expr + "\"\n"
@@ -691,10 +690,7 @@ CONTRACT_V3 = contract(["/stock", "/stock/{sku}", 
"/stock/{sku}/reserve"])
 REST_OPENAPI = "- restConfiguration:\n    apiContextPath: openapi\n\n- rest:\n 
   openApi:\n      specification: stock-api.json\n\n"
 REST_OPENAPI_VALIDATED = "- restConfiguration:\n    clientRequestValidation: 
true\n    apiContextPath: openapi\n\n- rest:\n    openApi:\n      
specification: stock-api.json\n\n"
 LIST_STOCK = route("listStock", DIRECT("listStock"), 
const("resource:file:stock.json"))
-LOOKUP = route("lookup", DIRECT("lookup"), const("resource:file:stock.json") + 
JSONPATH_SKU
-                + "        - choice:\n            when:\n              - 
expression:\n                  simple:\n                    expression: 
\"${body.size()} == 0\"\n"
-                  "                steps:\n                  - setBody:\n      
                expression:\n                        simple:\n                  
        expression: \"${null}\"\n"
-                  "            otherwise:\n              steps:\n              
  - setBody:\n                    expression:\n                      simple:\n  
                      expression: \"${body[0]}\"\n")
+LOOKUP = route("lookup", DIRECT("lookup"), const("resource:file:stock.json") + 
GROOVY_SKU)
 NULL_404 = ("        - choice:\n            when:\n              - 
expression:\n                  simple:\n                    expression: 
\"${body} == null\"\n"
             "                steps:\n" + set_header("CamelHttpResponseCode", 
"constant", "\"404\"", "                  ")
             + "                  - setBody:\n                      
expression:\n                        simple:\n                          
expression: '{\"error\": \"unknown sku ${header.sku}\"}'\n"
@@ -731,7 +727,7 @@ EXAMPLES.append({
     "seed": "contracts-openapi-server",
     "initial": {OS: os_s1, "stock-api.json": CONTRACT_V1, 
"application.properties": PORT_PROPS},
     "steps": [
-        {"request": "Add the operation getStock to the contract 
stock-api.json: GET /stock/{sku} with the path parameter sku, a 200 answering a 
StockItem and a 404 answering an Error. Then the route getStock (rest-openapi 
routes each operation to direct:<operationId>): it calls a helper route lookup 
that reads stock.json, picks the SKU with jsonpath $[?(@.sku == 
'${header.sku}')] (resultType java.util.List) and leaves the item as the body 
or null when the list is empty; getStock answers  [...]
+        {"request": "Add the operation getStock to the contract 
stock-api.json: GET /stock/{sku} with the path parameter sku, a 200 answering a 
StockItem and a 404 answering an Error. Then the route getStock (rest-openapi 
routes each operation to direct:<operationId>): it calls a helper route lookup 
that sets the body to the stock file (resource:file:stock.json, as listStock 
does), unmarshals the body with Jackson and picks the SKU with the Groovy 
expression body.find { it.sku == headers [...]
          "check": {"file_regex": "getStock", "files": {"stock-api.json": 
"\"operationId\": \"getStock\""}, "probes": [GET("/api/stock/CAMEL-MUG", 200, 
"\"sku\":\"CAMEL-MUG\",\"qty\":42"), GET("/api/stock/CAMEL-SOCKS", 404, 
"unknown sku CAMEL-SOCKS")]},
          "reference": {OS: os_s2, "stock-api.json": CONTRACT_V2}},
         {"request": "Add the operation reserveStock to the contract: POST 
/stock/{sku}/reserve with a JSON request body Reservation (orderId string, qty 
integer, both required), a 200 answering a ReservationResult (sku, reserved, 
remaining), and a 404. Then the route reserveStock: unmarshal the body with 
Jackson, keep it in an exchange property reservation, call direct:lookup, 
answer 404 for an unknown SKU as getStock does, otherwise log \"Reserved 
${exchangeProperty.reservation[qty]} x  [...]
diff --git a/ai-benchmark/ref-server.sh b/ai-benchmark/ref-server.sh
new file mode 100755
index 0000000..4018b03
--- /dev/null
+++ b/ai-benchmark/ref-server.sh
@@ -0,0 +1,17 @@
+#!/bin/zsh
+# Starts or stops the reference stock API server (the openapi-server example, 
copied to seed/openapi-server-ref by
+# gen_ladder.py) on port 8080 for the openapi-client example. Usage: 
ref-server.sh start|stop
+cd "$(dirname "$0")/seed/openapi-server-ref" || exit 2
+case "$1" in
+  start)
+    ( camel run * --logging-color=false > ../../ref-server.log 2>&1 ) &
+    for i in $(seq 1 60); do
+      curl -s -o /dev/null localhost:8080/api/openapi && { echo "ref server up 
after ${i}s"; exit 0; }
+      sleep 1
+    done
+    echo "ref server did not come up"; exit 1 ;;
+  stop)
+    # the server owns port 8080 by definition (jbang forks the java process, 
so a pattern on the command line misses it)
+    lsof -ti :8080 | xargs -r kill -TERM 2>/dev/null; sleep 3
+    lsof -ti :8080 | xargs -r kill -KILL 2>/dev/null; echo "ref server 
stopped" ;;
+esac
diff --git a/ai-benchmark/ref_pass.py b/ai-benchmark/ref_pass.py
new file mode 100755
index 0000000..cd3b089
--- /dev/null
+++ b/ai-benchmark/ref_pass.py
@@ -0,0 +1,36 @@
+#!/usr/bin/env python3
+"""Reference pass: runs every example of a set unchanged (the example's own 
files plus its seeds) through the same
+scorer the model is scored with. An example the reference cannot pass is a 
harness bug or an example bug, to be fixed
+before any model run. Usage: ref_pass.py [examples-ladder.json] [name ...]   
Results: ref/<name>/, ref.log"""
+import json, os, shutil, sys
+sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
+import agent_local as A
+
+REPO = os.path.expanduser(os.environ.get("BENCH_EXAMPLES_REPO", 
"~/workspace/camel-jbang-examples"))
+
+
+def main():
+    setfile = sys.argv[1] if len(sys.argv) > 1 else "examples-ladder.json"
+    only = set(sys.argv[2:])
+    examples = json.load(open(os.path.join(A.HERE, setfile)))
+    log = open(os.path.join(A.HERE, "ref.log"), "a")
+    passed = 0
+    for ex in examples:
+        if only and ex["name"] not in only:
+            continue
+        folder = os.path.join(A.HERE, "ref", ex["name"])
+        shutil.rmtree(folder, ignore_errors=True)
+        shutil.copytree(os.path.join(REPO, ex["example"]), folder, 
ignore=shutil.ignore_patterns("README.md", "metadata.json", "test", "parked"))
+        if ex.get("seed"):
+            A.seed_files(ex["name"], folder)
+        A.hook(ex.get("pre"), log)
+        ok, v, errs, nroutes, activity = A.run_folder(folder, 
ex["run_seconds"], ex.get("probe"), ex.get("probe_regex"), ex)
+        A.hook(ex.get("post"), log)
+        passed += ok
+        line = f"{ex['name']}: ok={ok} routes={nroutes} activity={activity}" + 
(("\n    " + errs.replace("\n", "\n    ")) if errs else "")
+        print(line, flush=True); print(line, file=log, flush=True)
+    print(f"reference pass: {passed} of {len(only) if only else len(examples)} 
pass")
+
+
+if __name__ == "__main__":
+    main()
diff --git a/ai-benchmark/run-http.sh b/ai-benchmark/run-http.sh
new file mode 100755
index 0000000..ae36e70
--- /dev/null
+++ b/ai-benchmark/run-http.sh
@@ -0,0 +1,8 @@
+#!/bin/zsh
+# Runs the four HTTP rungs of the stepwise ladder (CAMEL-24886): two servers 
checked with HTTP probes, a client calling a
+# server in the same app, and a contract-first client calling a peer app. A 
wrapper around run-stepwise-ladder.sh.
+# Usage: run-http.sh <tag> [k]                       Results: 
stepwise/<tag>-<i>/<name>/
+#        BENCH_REFERENCE=1 run-http.sh <tag>         the reference pass: the 
steps' own reference files, no model
+cd "$(dirname "$0")"
+export BENCH_ONLY="connect-stock-api connect-http-client 
contracts-openapi-server contracts-openapi-client"
+exec ./run-stepwise-ladder.sh "$@"
diff --git a/ai-benchmark/run-stepwise-ladder.sh 
b/ai-benchmark/run-stepwise-ladder.sh
index ae894aa..ae357b6 100755
--- a/ai-benchmark/run-stepwise-ladder.sh
+++ b/ai-benchmark/run-stepwise-ladder.sh
@@ -2,7 +2,7 @@
 # Runs the round-2 stepwise set: every steps-ladder/<name>.json, each from a 
fresh project, k times.
 # Usage: run-stepwise-ladder.sh <tag> [k]      Results: 
stepwise/<tag>-<i>/<name>/ (results.json, run.log, traces)
 # BENCH_REFERENCE=1 applies the reference files instead of asking the model 
(the reference pass of the steps files).
-# BENCH_ONLY=<name> runs one example.
+# BENCH_ONLY="<name> [<name> ...]" runs only those examples.
 set -u
 TAG="${1:?usage: run-stepwise-ladder.sh <tag> [k]}"
 K="${2:-1}"
@@ -14,7 +14,7 @@ for i in $(seq 1 "$K"); do
   : > "stepwise-$T.log"   # a fresh log per run, so a later wait on its DONE 
line cannot see an earlier run's
   for f in steps-ladder/*.json; do
     name=$(basename "$f" .json)
-    if [[ -n "${BENCH_ONLY:-}" && "$name" != "$BENCH_ONLY" ]]; then continue; 
fi
+    if [[ -n "${BENCH_ONLY:-}" && " $BENCH_ONLY " != *" $name "* ]]; then 
continue; fi
     python3 gen_stepwise.py "$name" > /dev/null
     echo "[$T] $name start $(date +%T)" | tee -a "stepwise-$T.log"
     BENCH_STEPS="$f" BENCH_TAG="$T/$name" python3 agent_mcp_stepwise.py > 
"stepwise-$T-$name.out" 2>&1
diff --git a/ai-benchmark/seed/_peer-openapi-server/openapi-server.camel.yaml 
b/ai-benchmark/seed/_peer-openapi-server/openapi-server.camel.yaml
index 23d364e..d2f497b 100644
--- a/ai-benchmark/seed/_peer-openapi-server/openapi-server.camel.yaml
+++ b/ai-benchmark/seed/_peer-openapi-server/openapi-server.camel.yaml
@@ -139,24 +139,11 @@
             expression:
               constant:
                 expression: resource:file:stock.json
+        - unmarshal:
+            json:
+              library: Jackson
+        # the item of that SKU as a map, or null when there is none
         - setBody:
             expression:
-              jsonpath:
-                expression: "$[?(@.sku == '${header.sku}')]"
-                resultType: java.util.List
-        - choice:
-            when:
-              - expression:
-                  simple:
-                    expression: "${body.size()} == 0"
-                steps:
-                  - setBody:
-                      expression:
-                        simple:
-                          expression: "${null}"
-            otherwise:
-              steps:
-                - setBody:
-                    expression:
-                      simple:
-                        expression: "${body[0]}"
+              groovy:
+                expression: "body.find { it.sku == headers.sku }"
diff --git a/ai-benchmark/summarize_stepwise.py 
b/ai-benchmark/summarize_stepwise.py
index fbac72f..8173c1b 100644
--- a/ai-benchmark/summarize_stepwise.py
+++ b/ai-benchmark/summarize_stepwise.py
@@ -1,12 +1,16 @@
 #!/usr/bin/env python3
-"""Per example and step, the passes over the runs of a stepwise ladder series: 
summarize_stepwise.py <tag> [k]"""
+"""Per example and step, the passes over the runs of a stepwise ladder series: 
summarize_stepwise.py <tag> [k]
+
+A step passes (ok) when its files, log and probes are right and it logged no 
new errors; final state right (ok_final)
+leaves the errors out: a model that tests its own app (the HTTP request tool) 
causes errors on the way to a right
+answer, which ok counts against it."""
 import json, os, sys
 tag = sys.argv[1]; k = int(sys.argv[2]) if len(sys.argv) > 2 else 1
 tags = [f"{tag}-{i}" for i in range(1, k + 1)] if k > 1 else [tag]
 names = sorted(os.path.splitext(f)[0] for f in os.listdir("steps-ladder") if 
f.endswith(".json"))
-print(f"| example | steps | passes per step ({' '.join(tags)}) | all steps 
passed |")
-print("|---|---|---|---|")
-total = 0; possible = 0; clean = 0
+print(f"| example | steps | passes per step ({' '.join(tags)}) | all steps 
passed | final state right |")
+print("|---|---|---|---|---|")
+total = 0; possible = 0; clean = 0; final = 0
 for n in names:
     per = []; runs = []
     for t in tags:
@@ -21,5 +25,6 @@ for n in names:
     for s in range(steps):
         ok = sum(1 for r in runs if s < len(r) and r[s]["ok"]); 
cols.append(f"{ok}/{len(runs)}"); total += ok; possible += len(runs)
     allok = sum(1 for r in runs if r and all(x["ok"] for x in r)); clean += 
allok
-    print(f"| {n} | {steps} | {' '.join(cols)} | {allok}/{len(runs)} |")
-print(f"\nsteps passed: {total} of {possible}; runs with every step passed: 
{clean}")
+    fin = sum(1 for r in runs for x in r if x.get("ok_final", x["ok"])); final 
+= fin
+    print(f"| {n} | {steps} | {' '.join(cols)} | {allok}/{len(runs)} | 
{fin}/{sum(len(r) for r in runs)} |")
+print(f"\nsteps passed: {total} of {possible}; final state right: {final} of 
{possible}; runs with every step passed: {clean}")

Reply via email to