{"model": "typesafe/jev-router", "task": "E4", "tier": "easy", "rep": 0, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.000291, "in_tok": 102, "out_tok": 217, "latency": 1.63, "id": "gen-1791297705-hAiSjc6STriQlqXdDJM4", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "typesafe/jev-router", "task": "E3", "tier": "easy", "rep": 0, "pass": true, "detail": "520", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.00073725, "in_tok": 78, "out_tok": 181, "latency": 3.13, "id": "gen-1791297704-EtOKwcANAcEHgyAvqyC3", "text": "<answer>520</answer>"}
{"model": "typesafe/jev-router", "task": "E2", "tier": "easy", "rep": 0, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "openai/gpt-6-luna", "provider": "Azure", "cost": 5.27e-05, "in_tok": 152, "out_tok": 75, "latency": 4.0, "id": "gen-1791297705-Pk4KyI1yObAHet3aABmh", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "typesafe/jev-router", "task": "E1", "tier": "easy", "rep": 0, "pass": true, "detail": "{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}", "routed": "openai/gpt-6-luna", "provider": "Azure", "cost": 4.11e-05, "in_tok": 136, "out_tok": 55, "latency": 4.11, "id": "gen-1791297705-p91nTlI7fxZhuJzcJJsI", "text": "<answer>{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}</answer>"}
{"model": "typesafe/jev-router", "task": "M2", "tier": "medium", "rep": 0, "pass": true, "detail": "37/37 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0020757, "in_tok": 151, "out_tok": 1692, "latency": 6.6, "id": "gen-1791297705-EJEyk0JRjDM0BbcrXTng", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    out = [0] * n\n    stack = []\n\n    for i in range(n - 1, -1, -1):\n        while stack and xs[stack[-1]] < xs[i]:\n            stack.pop()\n\n        if stack:\n            out[i] = xs[i] + (stack[-1] - i)\n        else:\n            out[i] = xs[i]\n\n        stack.append(i)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "typesafe/jev-router", "task": "E2", "tier": "easy", "rep": 1, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "openai/gpt-6-luna", "provider": "Azure", "cost": 4.87e-05, "in_tok": 152, "out_tok": 67, "latency": 3.19, "id": "gen-1791297709-Riughsvawkcgrh5o6u9f", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "typesafe/jev-router", "task": "E4", "tier": "easy", "rep": 1, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0002682, "in_tok": 102, "out_tok": 198, "latency": 1.21, "id": "gen-1791297712-iRNjN46MstvWtk0IoEEf", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "typesafe/jev-router", "task": "E3", "tier": "easy", "rep": 1, "pass": true, "detail": "520", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.000756, "in_tok": 78, "out_tok": 186, "latency": 3.46, "id": "gen-1791297711-9XOICPurzGbn29JWO6Zi", "text": "<answer>520</answer>"}
{"model": "typesafe/jev-router", "task": "M1", "tier": "medium", "rep": 0, "pass": true, "detail": "3, 32, 10, 8, 2, 9", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.004524, "in_tok": 272, "out_tok": 1152, "latency": 10.66, "id": "gen-1791297705-bnc86dsaiMWoDl3Eq6Ab", "text": "<answer>3, 32, 10, 8, 2, 9</answer>"}
{"model": "typesafe/jev-router", "task": "H4", "tier": "hard", "rep": 0, "pass": true, "detail": "grid ok", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.017694, "in_tok": 402, "out_tok": 1689, "latency": 12.63, "id": "gen-1791297705-xKWUMEw1lArBKVMwjPW1", "text": "I found this by case analysis on Dagny's position (Dagny lives in the teal house, clue 8).\n\n**Dagny's position**\n- Dagny in house 1 is impossible. Fenna would have to be in house 2, but house 2 drinks kvass and Fenna drinks chicha.\n- If Dagny is in house 3 or 4, Fenna is forced to the right of Dagny. Clue 5 then puts Bram in house 5, and Evander and Gideon take houses 1 and 2. Gideon must be in 2 with slate in 1 and ochre in 2, and Evander in 1 needs the mate drinker in house 2. House 2 drinks kvass, so both cases fail.\n- So Dagny is in house 2 (teal, kvass). Fenna is in house 1 or 3.\n\n**Fenna's position**\nTake Fenna in house 1 (chicha). Evander, Gideon and Bram fill houses 3 to 5.\n- Ochre cannot be house 2 (teal), so slate is in 3 and ochre in 4. Gideon must be right of slate, so Gideon is in 5.\n- The mate drinker is right of ochre, so mate is in house 5.\n- Evander must be next to house 5, so Evander is in 4 and Bram is in 3. Bram is slate, not ochre, so clue 6 holds.\n- The remaining colors, coral and olive, go to houses 1 and 5. Olive cannot be house 1, because sbiten would have to be immediately left of it. So olive is in 5 and coral in 1, with sbiten in house 4.\n- The last drink, horchata, goes to house 3.\n\nI did not check the case of Fenna in house 3, since this assignment satisfies every clue and the puzzle states the solution is unique.\n\n<answer>{\"1\": {\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"}, \"2\": {\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"}, \"3\": {\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"}, \"4\": {\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"}, \"5\": {\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "typesafe/jev-router", "task": "H2", "tier": "hard", "rep": 0, "pass": true, "detail": "46/46 hidden tests", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.005802, "in_tok": 221, "out_tok": 536, "latency": 11.58, "id": "gen-1791297706-iQyYgTYbkFEYGeRoQQBf", "text": "```python\ndef vault(grid: list[str]) -> int:\n    import heapq\n\n    if not grid or not grid[0]:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    start = None\n    portals = []\n\n    for r in range(rows):\n        for c in range(cols):\n            if grid[r][c] == \"S\":\n                start = (r, c)\n            elif grid[r][c] == \"*\":\n                portals.append((r, c))\n\n    if start is None:\n        return -1\n\n    portal_links = {}\n    if len(portals) == 2:\n        a, b = portals\n        portal_links[a] = b\n        portal_links[b] = a\n\n    sr, sc = start\n    distances = {(sr, sc, 0): 0}\n    heap = [(0, sr, sc, 0)]\n\n    while heap:\n        cost, r, c, keys = heapq.heappop(heap)\n        if cost != distances.get((r, c, keys)):\n            continue\n\n        if grid[r][c] == \"E\":\n            return cost\n\n        if (r, c) in portal_links:\n            nr, nc = portal_links[(r, c)]\n            state = (nr, nc, keys)\n            if cost < distances.get(state, float(\"inf\")):\n                distances[state] = cost\n                heapq.heappush(heap, (cost, nr, nc, keys))\n\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < rows and 0 <= nc < cols):\n                continue\n\n            cell = grid[nr][nc]\n            if cell == \"#\":\n                continue\n            if cell in \"ABC\" and not (keys & (1 << (ord(cell) - ord(\"A\")))):\n                continue\n\n            new_keys = keys\n            if cell in \"abc\":\n                new_keys |= 1 << (ord(cell) - ord(\"a\"))\n\n            new_cost = cost + 1\n            if cell in \"123456789\":\n                new_cost += int(cell)\n\n            state = (nr, nc, new_keys)\n            if new_cost < distances.get(state, float(\"inf\")):\n                distances[state] = new_cost\n                heapq.heappush(heap, (new_cost, nr, nc, new_keys))\n\n    return -1\n```"}
{"model": "typesafe/jev-router", "task": "M3", "tier": "medium", "rep": 0, "pass": true, "detail": "27 Dross 1130", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.006693, "in_tok": 164, "out_tok": 1752, "latency": 14.05, "id": "gen-1791297705-5FGtcJAR5vaL9VVEu0A7", "text": "<answer>27 Dross 1130</answer>"}
{"model": "typesafe/jev-router", "task": "M4", "tier": "medium", "rep": 0, "pass": true, "detail": "grid ok", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.00627225, "in_tok": 218, "out_tok": 1629, "latency": 14.2, "id": "gen-1791297704-afmAdHYVDtxvtAXXQjCd", "text": "<answer>\n{\n  \"1\": {\n    \"name\": \"Ilsa\",\n    \"color\": \"slate\",\n    \"drink\": \"mate\"\n  },\n  \"2\": {\n    \"name\": \"Bram\",\n    \"color\": \"plum\",\n    \"drink\": \"horchata\"\n  },\n  \"3\": {\n    \"name\": \"Dagny\",\n    \"color\": \"olive\",\n    \"drink\": \"chicha\"\n  },\n  \"4\": {\n    \"name\": \"Evander\",\n    \"color\": \"teal\",\n    \"drink\": \"kvass\"\n  }\n}\n</answer>"}
{"model": "typesafe/jev-router", "task": "H3", "tier": "hard", "rep": 0, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.014864, "in_tok": 292, "out_tok": 1428, "latency": 11.94, "id": "gen-1791297708-mRpL8kVxlzvaQCpTmj9S", "text": "This schedule finishes at time 16. I did not run a search; the argument below shows 16 is the best possible.\n\n**Lower bound.** J8 is hot, so it cannot overlap with J2 or J5. Every other job is a predecessor of J8, so none of them can run beside it either. Thus the second machine sits idle while J8 runs its 5 units.\n\nBefore J8 starts, all of its predecessors must finish. J2 must also fit somewhere. J2 cannot run beside J8, so it either finishes before J8 starts or runs alone after J8.\n\n- If J2 runs before J8, the other 21 units of work (everything except J8) must fit on 2 machines before J8 starts. That forces J8 to start at 11 or later, so the makespan is at least 11 + 5 = 16.\n- If J2 runs after J8, J8 can start at 10 at the earliest. The makespan is then at least 10 + 5 + 2 = 17.\n\nSo the makespan is at least 16.\n\n**Schedule achieving 16.**\n\n- Machine A: J3 runs 0\u20131, J6 runs 1\u20137, J5 runs 7\u20139, J2 runs 9\u201311, J8 runs 11\u201316.\n- Machine B: J1 runs 0\u20135, J4 runs 5\u20136, J7 runs 6\u201310, then idle.\n\nPrecedence checks:\n- J3 finishes at 1, before J6 (starts 1) and J8 (starts 11).\n- J1 finishes at 5, before J4 (starts 5).\n- J4 finishes at 6, before J8.\n- J5 finishes at 9, before J8.\n- J6 finishes at 7, before J8.\n- J7 finishes at 10, before J8.\n\nHeat check: J5 runs 7\u20139, J2 runs 9\u201311 and J8 runs 11\u201316, so no two hot jobs overlap.\n\n<answer>{\"J1\":0,\"J2\":9,\"J3\":0,\"J4\":5,\"J5\":7,\"J6\":1,\"J7\":6,\"J8\":11}</answer>"}
{"model": "typesafe/jev-router", "task": "E1", "tier": "easy", "rep": 1, "pass": true, "detail": "{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}", "routed": "openai/gpt-6-luna", "provider": "Azure", "cost": 3.96e-05, "in_tok": 136, "out_tok": 52, "latency": 11.5, "id": "gen-1791297708-WhRPheFdKDI7JyP1fIss", "text": "<answer>{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}</answer>"}
{"model": "typesafe/jev-router", "task": "M2", "tier": "medium", "rep": 1, "pass": true, "detail": "37/37 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0018981, "in_tok": 151, "out_tok": 1544, "latency": 5.2, "id": "gen-1791297715-WSSwoQDHzR7NvPSGy6dc", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    out = [0] * n\n    stack = []\n\n    for i in range(n - 1, -1, -1):\n        x = xs[i]\n        while stack and xs[stack[-1]] < x:\n            stack.pop()\n\n        if stack:\n            out[i] = x + (stack[-1] - i)\n        else:\n            out[i] = x\n\n        stack.append(i)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "typesafe/jev-router", "task": "H1", "tier": "hard", "rep": 0, "pass": false, "detail": "3, 87846, 4, -586, -573", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.024566, "in_tok": 453, "out_tok": 2366, "latency": 16.27, "id": "gen-1791297705-3YpDG1F7VrJdqYlzZA93", "text": "I traced the program token by token, tracking the full stack after each step.\n\n1. `8 3 6` gives [8, 3, 6]. `rep 4 [11 mul]` multiplies the top by 11 four times, so 6 becomes 6\u00b714641 = 87846. The stack is [8, 3, 87846].\n2. `rot` gives [3, 87846, 8].\n3. `rep 4 [10 add 4 5]` runs four times, each time adding 10 to the top and then pushing 4 and 5. The stack becomes [3, 87846, 18, 4, 15, 4, 15, 4, 15, 4, 5].\n4. `over` pushes 4, giving [..., 15, 4, 5, 4].\n5. `rep 3 [drop zig]`:\n   - Iteration 1 drops 4, then zig turns 5 into 16.\n   - Iteration 2 drops 16, then zig turns 4 into 2.\n   - Iteration 3 drops 2, then zig turns 15 into 46.\n   - The stack is [3, 87846, 18, 4, 15, 4, 15, 4, 46].\n6. `6 add` makes the top 52. `over` pushes 4, and `3` pushes 3. The stack is [..., 15, 4, 52, 4, 3].\n7. `rep 2 [mul 2 drop]`:\n   - Iteration 1: mul gives 4\u00b73 = 12, and `2 drop` is a no-op.\n   - Iteration 2: mul gives 52\u00b712 = 624, and `2 drop` is again a no-op.\n   - The stack is [3, 87846, 18, 4, 15, 4, 15, 4, 624].\n8. `sub` gives 4 \u2212 624 = \u2212620, so the stack is [..., 15, 4, 15, \u2212620].\n9. `2 ifz [...]` pops 2, which is nonzero, so the body is skipped.\n10. `add` gives 15 + (\u2212620) = \u2212605, so the stack is [3, 87846, 18, 4, 15, 4, \u2212605].\n11. `mod` gives 4 mod \u2212605 = \u2212601, so the stack is [3, 87846, 18, 4, 15, \u2212601].\n12. `mod` gives 15 mod \u2212601 = \u2212586, so the stack is [3, 87846, 18, 4, \u2212586].\n13. `rot` turns (18, 4, \u2212586) into (4, \u2212586, 18), giving [3, 87846, 4, \u2212586, 18].\n14. `5 sub` gives 13, so the stack is [3, 87846, 4, \u2212586, 13].\n15. `over` pushes \u2212586. `mod` gives 13 mod \u2212586 = \u2212573, so the stack is [3, 87846, 4, \u2212586, \u2212573].\n16. `1 ifz [...]` pops 1, which is nonzero, so the body is skipped.\n\n<answer>3, 87846, 4, -586, -573</answer>"}
{"model": "typesafe/jev-router", "task": "M1", "tier": "medium", "rep": 1, "pass": true, "detail": "3, 32, 10, 8, 2, 9", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.0046365, "in_tok": 272, "out_tok": 1182, "latency": 9.16, "id": "gen-1791297713-sS7qDq51B0w6XW617DGo", "text": "<answer>3, 32, 10, 8, 2, 9</answer>"}
{"model": "typesafe/jev-router", "task": "E4", "tier": "easy", "rep": 2, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.000267, "in_tok": 102, "out_tok": 197, "latency": 1.14, "id": "gen-1791297722-pUNLJEId5fjHsa9ivTBn", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "typesafe/jev-router", "task": "E1", "tier": "easy", "rep": 2, "pass": true, "detail": "{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}", "routed": "openai/gpt-6-luna", "provider": "Azure", "cost": 3.91e-05, "in_tok": 136, "out_tok": 51, "latency": 3.95, "id": "gen-1791297720-wi2Wfv8YoG2txil4jobm", "text": "<answer>{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}</answer>"}
{"model": "typesafe/jev-router", "task": "E3", "tier": "easy", "rep": 2, "pass": true, "detail": "520", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.00065475, "in_tok": 78, "out_tok": 159, "latency": 3.52, "id": "gen-1791297721-A61jvCpyQnVMgskyQcrF", "text": "<answer>520</answer>"}
{"model": "typesafe/jev-router", "task": "E2", "tier": "easy", "rep": 2, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "openai/gpt-6-luna", "provider": "Azure", "cost": 5.07e-05, "in_tok": 152, "out_tok": 71, "latency": 4.65, "id": "gen-1791297721-S83hhLA04Uu0sN3F4iOJ", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "typesafe/jev-router", "task": "H1", "tier": "hard", "rep": 1, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.022826, "in_tok": 453, "out_tok": 2192, "latency": 13.91, "id": "gen-1791297715-dyHVk1eE2hS30FasNzd0", "text": "Simulating the program step by step (stack shown bottom to top):\n\n1. `8 3 6` gives [8, 3, 6].\n2. `rep 4 [11 mul]` multiplies the top by 11 four times: 6\u00b714641 = 87846, so the stack is [8, 3, 87846].\n3. `rot` gives [3, 87846, 8].\n4. `10` gives [3, 87846, 8, 10].\n5. `rep 4 [10 add 4 5]` leaves [3, 87846, 8, 20, 4, 15, 4, 15, 4, 15, 4, 5].\n6. `over` gives [..., 4, 5, 4].\n7. `rep 3 [drop zig]`:\n   - Pass 1 drops 4 and zig turns 5 into 16.\n   - Pass 2 drops 16 and zig turns 4 into 2.\n   - Pass 3 drops 2 and zig turns 15 into 46.\n   - The stack is [3, 87846, 8, 20, 4, 15, 4, 15, 4, 46].\n8. `6 add` makes the top 52.\n9. `over` then `3` gives [..., 15, 4, 52, 4, 3].\n10. `rep 2 [mul 2 drop]`: 4\u00b73 = 12, then 52\u00b712 = 624. The stack is [3, 87846, 8, 20, 4, 15, 4, 15, 4, 624].\n11. `sub`: 4 \u2212 624 = \u2212620, giving [3, 87846, 8, 20, 4, 15, 4, 15, \u2212620].\n12. `2 ifz [...]` pops 2, which is nonzero, so the body is skipped.\n13. `add`: 15 + (\u2212620) = \u2212605, giving [3, 87846, 8, 20, 4, 15, 4, \u2212605].\n14. `mod`: 4 mod \u2212605 = \u2212601, giving [..., 4, 15, \u2212601].\n15. `mod`: 15 mod \u2212601 = \u2212586, giving [3, 87846, 8, 20, 4, \u2212586].\n16. `rot` gives [3, 87846, 8, 4, \u2212586, 20].\n17. `5 sub`: 20 \u2212 5 = 15, giving [3, 87846, 8, 4, \u2212586, 15].\n18. `over` gives [..., \u2212586, 15, \u2212586].\n19. `mod`: 15 mod \u2212586 = \u2212571, giving [3, 87846, 8, 4, \u2212586, \u2212571].\n20. `1 ifz [...]` pops 1, which is nonzero, so the body is skipped.\n\n<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "typesafe/jev-router", "task": "M3", "tier": "medium", "rep": 1, "pass": true, "detail": "27 Dross 1130", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.0064455, "in_tok": 164, "out_tok": 1686, "latency": 12.4, "id": "gen-1791297717-h2w2Sja5V7h1i0TOVPU9", "text": "<answer>27 Dross 1130</answer>"}
{"model": "typesafe/jev-router", "task": "H2", "tier": "hard", "rep": 1, "pass": true, "detail": "46/46 hidden tests", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.005892, "in_tok": 221, "out_tok": 545, "latency": 12.37, "id": "gen-1791297719-3G29P2PxJETqLhuS6l24", "text": "```python\ndef vault(grid: list[str]) -> int:\n    import heapq\n\n    if not grid or not grid[0]:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    start = end = None\n    portals = []\n\n    for r, row in enumerate(grid):\n        for c, cell in enumerate(row):\n            if cell == \"S\":\n                start = (r, c)\n            elif cell == \"E\":\n                end = (r, c)\n            elif cell == \"*\":\n                portals.append((r, c))\n\n    if start is None or end is None:\n        return -1\n\n    links = {}\n    if len(portals) == 2:\n        a, b = portals\n        links[a] = b\n        links[b] = a\n\n    initial = (start[0], start[1], 0)\n    distances = {initial: 0}\n    heap = [(0, *initial)]\n\n    while heap:\n        cost, r, c, keys = heapq.heappop(heap)\n        if cost != distances.get((r, c, keys)):\n            continue\n        if (r, c) == end:\n            return cost\n\n        if (r, c) in links:\n            nr, nc = links[(r, c)]\n            state = (nr, nc, keys)\n            if cost < distances.get(state, float(\"inf\")):\n                distances[state] = cost\n                heapq.heappush(heap, (cost, nr, nc, keys))\n\n        for dr, dc in ((1, 0), (-1, 0), (0, 1), (0, -1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < rows and 0 <= nc < cols):\n                continue\n\n            cell = grid[nr][nc]\n            if cell == \"#\":\n                continue\n            if cell in \"ABC\" and not (keys & (1 << (ord(cell) - ord(\"A\")))):\n                continue\n\n            next_keys = keys\n            if cell in \"abc\":\n                next_keys |= 1 << (ord(cell) - ord(\"a\"))\n\n            next_cost = cost + 1\n            if cell in \"123456789\":\n                next_cost += int(cell)\n\n            state = (nr, nc, next_keys)\n            if next_cost < distances.get(state, float(\"inf\")):\n                distances[state] = next_cost\n                heapq.heappush(heap, (next_cost, nr, nc, next_keys))\n\n    return -1\n```"}
{"model": "typesafe/jev-router", "task": "M2", "tier": "medium", "rep": 2, "pass": true, "detail": "37/37 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0023793, "in_tok": 151, "out_tok": 1945, "latency": 7.03, "id": "gen-1791297724-nJuunGkLT2uEPOtkxzFN", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    out = [0] * n\n    stack = []\n\n    for i in range(n - 1, -1, -1):\n        while stack and xs[stack[-1]] < xs[i]:\n            stack.pop()\n\n        if stack:\n            out[i] = xs[i] + (stack[-1] - i)\n        else:\n            out[i] = xs[i]\n\n        stack.append(i)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "typesafe/jev-router", "task": "M4", "tier": "medium", "rep": 1, "pass": true, "detail": "grid ok", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.0055785, "in_tok": 218, "out_tok": 1444, "latency": 14.43, "id": "gen-1791297718-ldELjXCdEATnDEmfOUdQ", "text": "<answer>{\"1\": {\"name\": \"Ilsa\", \"color\": \"slate\", \"drink\": \"mate\"}, \"2\": {\"name\": \"Bram\", \"color\": \"plum\", \"drink\": \"horchata\"}, \"3\": {\"name\": \"Dagny\", \"color\": \"olive\", \"drink\": \"chicha\"}, \"4\": {\"name\": \"Evander\", \"color\": \"teal\", \"drink\": \"kvass\"}}</answer>"}
{"model": "typesafe/jev-router", "task": "H3", "tier": "hard", "rep": 1, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.014654, "in_tok": 292, "out_tok": 1407, "latency": 13.04, "id": "gen-1791297720-tEYXTEe7HmF4KhRMi8E8", "text": "This schedule finishes at time 16; I argued that 16 is the minimum but did not check every alternative exhaustively.\n\n**Lower bound.** J8 is hot and every other job except J2 must finish before J8 starts. J2 is also hot, so it cannot overlap J8 either.\n\n- If J2 runs before J8, then 21 units of work (everything except J8) must fit before J8 starts at time s. On 2 machines this needs s \u2265 10.5, so s \u2265 11 and the makespan is at least 16.\n- If J2 runs after J8, then 19 units of work must fit before s, so s \u2265 10. J8 ends at 15 and J2 then runs 15\u201317, giving 17, which is worse.\n\nSo 16 is the best possible.\n\n**Schedule achieving 16.**\n- Machine A: J1 runs 0\u20135, J4 runs 5\u20136, J7 runs 6\u201310, then the machine is idle until J8.\n- Machine B: J3 runs 0\u20131, J6 runs 1\u20137, J5 runs 7\u20139, J2 runs 9\u201311.\n- J8 runs 11\u201316 on either machine.\n\n**Checks.**\n- Precedences: J4 starts at 5, after J1 finishes at 5. J6 starts at 1, after J3 finishes at 1. J8 starts at 11, after all of J3, J4, J5, J6 and J7 have finished (the last of these, J2 aside, ends at 10).\n- Heat rule: J5 runs 7\u20139, J2 runs 9\u201311 and J8 runs 11\u201316, so no two hot jobs overlap.\n\n<answer>{\"J1\":0,\"J2\":9,\"J3\":0,\"J4\":5,\"J5\":7,\"J6\":1,\"J7\":6,\"J8\":11}</answer>"}
{"model": "typesafe/jev-router", "task": "M1", "tier": "medium", "rep": 2, "pass": true, "detail": "3, 32, 10, 8, 2, 9", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.0047715, "in_tok": 272, "out_tok": 1218, "latency": 11.52, "id": "gen-1791297723-aMbu1gTUGayRt4Gz6OvI", "text": "<answer>3, 32, 10, 8, 2, 9</answer>"}
{"model": "nvidia/switchyard", "task": "E1", "tier": "easy", "rep": 0, "pass": true, "detail": "{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.002256, "in_tok": 216, "out_tok": 100, "latency": 2.55, "id": "gen-1791297732-9NTnmBHO1EDd8C5HmeOG", "text": "<answer>{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}</answer>"}
{"model": "nvidia/switchyard", "task": "E2", "tier": "easy", "rep": 0, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 0.000119286, "in_tok": 178, "out_tok": 167, "latency": 2.68, "id": "gen-1791297733-VJNn2ERsKEnRqHnFmftP", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "typesafe/jev-router", "task": "H4", "tier": "hard", "rep": 1, "pass": true, "detail": "grid ok", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.025344, "in_tok": 402, "out_tok": 2454, "latency": 17.57, "id": "gen-1791297719-US9dlxBHPMhVxSB8YB3e", "text": "This assignment satisfies all eleven clues. I found it by case analysis on where the sbiten drinker sits, which fixes the olive house just to its right.\n\n1. By clue 4, house 2 drinks kvass, so sbiten (clue 3) can only be in house 1, 3 or 4, with the olive house at 2, 4 or 5 respectively.\n2. Fenna drinks chicha (clue 9), so she cannot be in house 2. She must also be next to Dagny (clue 10), who is the teal house (clue 8).\n3. Sbiten in house 1 (olive 2) forces Dagny into house 1 and Fenna into house 2. That contradicts clue 4.\n4. Sbiten in house 3 (olive 4) fails in every placement of the slate/ochre pair. Each placement either forces Fenna into house 2 again or leaves no valid seat for Evander next to the mate drinker.\n5. Sbiten in house 4 (olive 5) works only when slate is in house 3 and ochre in house 4:\n   - Ochre in house 4 puts the mate drinker in house 5.\n   - Teal and coral occupy houses 1 and 2. Dagny in house 1 would force Fenna into house 2, so Dagny is teal in house 2.\n   - Evander must be next to the mate drinker in house 5, so Evander is in house 4.\n   - Gideon must be right of slate (house 3), so Gideon is in house 5.\n   - Bram is right of Dagny and not ochre (house 4), so Bram is in house 3. Fenna takes house 1.\n   - Fenna drinks chicha, so house 1 is chicha. House 3 is left with horchata.\n   - House 1 is coral, the one color left.\n\n<answer>{\"1\": {\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\n\"2\": {\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\n\"3\": {\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\n\"4\": {\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\n\"5\": {\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "nvidia/switchyard", "task": "E4", "tier": "easy", "rep": 0, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.000267, "in_tok": 102, "out_tok": 197, "latency": 1.72, "id": "gen-1791297735-nxyqLQC6EvjJBistDhge", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "nvidia/switchyard", "task": "E3", "tier": "easy", "rep": 0, "pass": true, "detail": "520", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0027765, "in_tok": 159, "out_tok": 148, "latency": 2.15, "id": "gen-1791297735-8g86T2bNi0lXWEKiFHj1", "text": "<answer>520</answer>"}
{"model": "typesafe/jev-router", "task": "H1", "tier": "hard", "rep": 2, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.021646, "in_tok": 453, "out_tok": 2074, "latency": 13.87, "id": "gen-1791297724-4bCddX05V2wcFEg72Hz1", "text": "Final stack, bottom to top: 3, 87846, 8, 4, -586, -571. I traced the program step by step, tracking the full stack after each token.\n\n1. `8 3 6` gives [8, 3, 6]. `rep 4 [11 mul]` multiplies the top by 11 four times: 6\u00b714641 = 87846, so the stack is [8, 3, 87846].\n2. `rot` gives [3, 87846, 8]. Then `10` gives [3, 87846, 8, 10].\n3. `rep 4 [10 add 4 5]` leaves [3, 87846, 8, 20, 4, 15, 4, 15, 4, 15, 4, 5].\n4. `over` pushes a copy of the second-from-top, giving \u2026, 4, 5, 4.\n5. `rep 3 [drop zig]`:\n   - drop 4, then zig 5 gives 16.\n   - drop 16, then zig 4 gives 2.\n   - drop 2, then zig 15 gives 46.\n   \n   The stack is [3, 87846, 8, 20, 4, 15, 4, 15, 4, 46].\n6. `6 add` makes the top 52. `over` copies 4 and `3` is pushed, giving \u2026, 4, 52, 4, 3.\n7. `rep 2 [mul 2 drop]`: 4\u00b73 = 12, then 52\u00b712 = 624. The stack is [3, 87846, 8, 20, 4, 15, 4, 15, 4, 624].\n8. `sub` gives 4 \u2212 624 = \u2212620. Then `2 ifz [...]` pops 2, which is nonzero, so the body is skipped. The stack is [3, 87846, 8, 20, 4, 15, 4, 15, \u2212620].\n9. `add` gives 15 + (\u2212620) = \u2212605, so the stack is [3, 87846, 8, 20, 4, 15, 4, \u2212605].\n10. `mod` gives 4 mod \u2212605 = \u2212601. The next `mod` gives 15 mod \u2212601 = \u2212586. The stack is [3, 87846, 8, 20, 4, \u2212586].\n11. `rot` gives [3, 87846, 8, 4, \u2212586, 20]. Then `5 sub` makes the top 15.\n12. `over` copies \u2212586, and `mod` gives 15 mod \u2212586 = \u2212571. The stack is [3, 87846, 8, 4, \u2212586, \u2212571].\n13. `1 ifz [...]` pops 1, which is nonzero, so the body is skipped.\n\n<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "typesafe/jev-router", "task": "M4", "tier": "medium", "rep": 2, "pass": true, "detail": "grid ok", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.00491475, "in_tok": 218, "out_tok": 1267, "latency": 9.83, "id": "gen-1791297729-HLtdcRkOJ0VrCY1bZVKh", "text": "<answer>{\"1\": {\"name\": \"Ilsa\", \"color\": \"slate\", \"drink\": \"mate\"}, \"2\": {\"name\": \"Bram\", \"color\": \"plum\", \"drink\": \"horchata\"}, \"3\": {\"name\": \"Dagny\", \"color\": \"olive\", \"drink\": \"chicha\"}, \"4\": {\"name\": \"Evander\", \"color\": \"teal\", \"drink\": \"kvass\"}}</answer>"}
{"model": "nvidia/switchyard", "task": "M1", "tier": "medium", "rep": 0, "pass": true, "detail": "3,32,10,8,2,9", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 0.000658047, "in_tok": 291, "out_tok": 1094, "latency": 6.48, "id": "gen-1791297735-bYXzQYeBFZBwcCFkqqTs", "text": "<answer>3,32,10,8,2,9</answer>"}
{"model": "typesafe/jev-router", "task": "M3", "tier": "medium", "rep": 2, "pass": true, "detail": "27 Dross 1130", "routed": "google/gemini-3.8-flash", "provider": "Google AI Studio", "cost": 0.009483, "in_tok": 164, "out_tok": 2496, "latency": 17.72, "id": "gen-1791297725-5XGbZLtZyoGSru64RvmU", "text": "<answer>27 Dross 1130</answer>"}
{"model": "typesafe/jev-router", "task": "H4", "tier": "hard", "rep": 2, "pass": true, "detail": "grid ok", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.024234, "in_tok": 402, "out_tok": 2343, "latency": 17.27, "id": "gen-1791297730-mQB3ndi5EPfJy6lNwJ0c", "text": "Houses 1 to 5 hold the following residents, colors and drinks.\n\n| House | Name | Color | Drink |\n|---|---|---|---|\n| 1 | Fenna | coral | chicha |\n| 2 | Dagny | teal | kvass |\n| 3 | Bram | slate | horchata |\n| 4 | Evander | ochre | sbiten |\n| 5 | Gideon | olive | mate |\n\nI found it by case analysis on where the slate and ochre houses (which must be adjacent) can sit.\n\n- **Slate/ochre in houses 1\u20132, 2\u20131, 3\u20132 or 2\u20133:** every one of these placements forces a contradiction. Dagny (teal) ends up in house 3 or 4, or in house 1. When she is in house 3 or 4, Fenna (chicha) must be next to her, but the neighboring houses are taken by kvass (house 2), sbiten, or Bram. When she is in house 1, Fenna would have to be in house 2, which drinks kvass.\n- **Slate in 3, ochre in 4:** the mate drinker must be in house 5, so Evander is in house 4. Gideon must be right of slate, so he is in house 5. Dagny cannot be in house 1, because Fenna would then be in house 2, which drinks kvass. So Dagny is in house 2 (teal), Fenna in house 1 (chicha), and Bram in house 3.\n- **Remaining colors and drinks:** olive must come right after sbiten, so olive is in house 5, sbiten in house 4, and coral is left for house 1. Horchata is left for house 3.\n\nAll eleven clues check against this arrangement.\n\n<answer>{\"1\": {\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"}, \"2\": {\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"}, \"3\": {\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"}, \"4\": {\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"}, \"5\": {\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "nvidia/switchyard", "task": "E1", "tier": "easy", "rep": 1, "pass": true, "detail": "{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 7.8678e-05, "in_tok": 158, "out_tok": 100, "latency": 2.32, "id": "gen-1791297747-ECmbfd2NUeAKqcJAZ6nA", "text": "<answer>{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}</answer>"}
{"model": "nvidia/switchyard", "task": "M2", "tier": "medium", "rep": 0, "pass": true, "detail": "37/37 hidden tests", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0218985, "in_tok": 201, "out_tok": 1413, "latency": 12.76, "id": "gen-1791297737-xWgu4H7DlWhR6aUaAuj0", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    dist = [0] * n\n    stack: list[int] = []\n\n    for i in range(n - 1, -1, -1):\n        while stack and xs[stack[-1]] < xs[i]:\n            stack.pop()\n        if stack:\n            dist[i] = stack[-1] - i\n        stack.append(i)\n\n    out = [xs[i] + dist[i] for i in range(n)]\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "nvidia/switchyard", "task": "M4", "tier": "medium", "rep": 0, "pass": true, "detail": "grid ok", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.024563, "in_tok": 298, "out_tok": 1568, "latency": 12.5, "id": "gen-1791297738-HiAj4jnOhLd6mQxmHC0J", "text": "<answer>{\"1\":{\"name\":\"Ilsa\",\"color\":\"slate\",\"drink\":\"mate\"},\"2\":{\"name\":\"Bram\",\"color\":\"plum\",\"drink\":\"horchata\"},\"3\":{\"name\":\"Dagny\",\"color\":\"olive\",\"drink\":\"chicha\"},\"4\":{\"name\":\"Evander\",\"color\":\"teal\",\"drink\":\"kvass\"}}</answer>"}
{"model": "typesafe/jev-router", "task": "H3", "tier": "hard", "rep": 2, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.024314, "in_tok": 292, "out_tok": 2373, "latency": 20.3, "id": "gen-1791297731-Wh9ZGlvUmFMrnRkrKQWv", "text": "This schedule gives a makespan of 16. I checked it against all constraints, and a counting argument shows 15 or less is not achievable, though I did not enumerate every schedule exhaustively.\n\n**Lower bound.** The total work is 5+2+1+1+2+6+4+5 = 26, so two machines need at least 13. J8 must wait for J3, J4, J5, J6 and J7. The hot rule also means J2 cannot overlap J8: it must finish before J8 starts or start after J8 ends. Let s be J8's start time.\n\n- The other seven jobs total 21 units of work.\n- If J2 runs before J8, all seven other jobs finish by s, so they need 2s \u2265 21, which gives s \u2265 11. J8 then ends at 16 or later.\n- If J2 runs after J8, it ends at s+7 or later. The six predecessors of J8 (19 units of work) must fit in 2s units, so s \u2265 10, and the makespan is then at least 17.\n- Either way the makespan is at least 16, so 15 or less is impossible.\n\n**Schedule achieving 16.**\n- Machine 1: J1 runs 0\u20135, J4 runs 5\u20136, J2 runs 6\u20138, J5 runs 8\u201310, and J8 runs 11\u201316.\n- Machine 2: J3 runs 0\u20131, J6 runs 1\u20137, and J7 runs 7\u201311.\n\nPrecedence holds: J4 starts at 5 when J1 finishes, J6 starts at 1 when J3 finishes, and J8 starts at 11 after J7 finishes. J2 (6\u20138), J5 (8\u201310) and J8 (11\u201316) never overlap, so the heat rule holds.\n\n<answer>{\"J1\":0,\"J2\":6,\"J3\":0,\"J4\":5,\"J5\":8,\"J6\":1,\"J7\":7,\"J8\":11}</answer>"}
{"model": "nvidia/switchyard", "task": "E2", "tier": "easy", "rep": 1, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 0.000143538, "in_tok": 178, "out_tok": 210, "latency": 3.07, "id": "gen-1791297749-hPdSYzxJv0DK0kZp0Ds4", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "nvidia/switchyard", "task": "E4", "tier": "easy", "rep": 1, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "deepseek/deepseek-v4.1-flash", "provider": "InferenceNet", "cost": 7.704e-05, "in_tok": 102, "out_tok": 233, "latency": 2.04, "id": "gen-1791297750-b2LjCQM3nq8GDiCdO474", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "nvidia/switchyard", "task": "E3", "tier": "easy", "rep": 1, "pass": true, "detail": "520", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0033015, "in_tok": 159, "out_tok": 183, "latency": 3.35, "id": "gen-1791297750-Ok3HccAPijrG48YulAIq", "text": "<answer>520</answer>"}
{"model": "nvidia/switchyard", "task": "M3", "tier": "medium", "rep": 0, "pass": true, "detail": "27 Dross 1130", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0381595, "in_tok": 227, "out_tok": 2491, "latency": 22.16, "id": "gen-1791297737-UZciRhMcPXOOMYKItosw", "text": "<answer>27 Dross 1130</answer>"}
{"model": "typesafe/jev-router", "task": "H2", "tier": "hard", "rep": 2, "pass": true, "detail": "46/46 hidden tests", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.009162, "in_tok": 221, "out_tok": 872, "latency": 27.84, "id": "gen-1791297731-tCLQ9qEEkZRVMS7rXoEV", "text": "```python\ndef vault(grid: list[str]) -> int:\n    from heapq import heappop, heappush\n\n    if not grid or not grid[0]:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    cells = \"\".join(grid)\n    start, end = cells.find(\"S\"), cells.find(\"E\")\n    if start == -1 or end == -1:\n        return -1\n\n    portals = [i for i, cell in enumerate(cells) if cell == \"*\"]\n    linked = {}\n    if len(portals) == 2:\n        a, b = portals\n        linked = {a: b, b: a}\n\n    distances = [[float(\"inf\")] * 8 for _ in cells]\n    distances[start][0] = 0\n    heap = [(0, start, 0)]\n\n    while heap:\n        cost, pos, keys = heappop(heap)\n        if cost != distances[pos][keys]:\n            continue\n        if pos == end:\n            return cost\n\n        destination = linked.get(pos)\n        if destination is not None and cost < distances[destination][keys]:\n            distances[destination][keys] = cost\n            heappush(heap, (cost, destination, keys))\n\n        row, col = divmod(pos, cols)\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = row + dr, col + dc\n            if not (0 <= nr < rows and 0 <= nc < cols):\n                continue\n\n            nxt = nr * cols + nc\n            cell = cells[nxt]\n            if cell == \"#\":\n                continue\n            if \"A\" <= cell <= \"C\":\n                required = 1 << (ord(cell) - ord(\"A\"))\n                if not (keys & required):\n                    continue\n\n            next_keys = keys\n            if \"a\" <= cell <= \"c\":\n                next_keys |= 1 << (ord(cell) - ord(\"a\"))\n\n            step_cost = 1\n            if \"1\" <= cell <= \"9\":\n                step_cost += int(cell)\n\n            next_cost = cost + step_cost\n            if next_cost < distances[nxt][next_keys]:\n                distances[nxt][next_keys] = next_cost\n                heappush(heap, (next_cost, nxt, next_keys))\n\n    return -1\n```"}
{"model": "nvidia/switchyard", "task": "M2", "tier": "medium", "rep": 1, "pass": true, "detail": "37/37 hidden tests", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0124035, "in_tok": 201, "out_tok": 780, "latency": 7.89, "id": "gen-1791297752-qE8IXnZdOKQcUAo5pfSU", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n    out = [0] * n\n    stack: list[int] = []  # indices with xs values in non-decreasing order from top\n    for i in range(n - 1, -1, -1):\n        xi = xs[i]\n        while stack and xs[stack[-1]] < xi:\n            stack.pop()\n        d = stack[-1] - i if stack else 0\n        out[i] = xi + d\n        stack.append(i)\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "nvidia/switchyard", "task": "M1", "tier": "medium", "rep": 1, "pass": true, "detail": "3,32,10,8,2,9", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0239255, "in_tok": 343, "out_tok": 1515, "latency": 11.43, "id": "gen-1791297752-Nea84GBOsdKfCPPwqlvx", "text": "<answer>3,32,10,8,2,9</answer>"}
{"model": "nvidia/switchyard", "task": "M4", "tier": "medium", "rep": 1, "pass": true, "detail": "grid ok", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.022298, "in_tok": 298, "out_tok": 1417, "latency": 13.51, "id": "gen-1791297759-XwiBwK1rmbIYcLixikNC", "text": "<answer>{\"1\":{\"name\":\"Ilsa\",\"color\":\"slate\",\"drink\":\"mate\"},\"2\":{\"name\":\"Bram\",\"color\":\"plum\",\"drink\":\"horchata\"},\"3\":{\"name\":\"Dagny\",\"color\":\"olive\",\"drink\":\"chicha\"},\"4\":{\"name\":\"Evander\",\"color\":\"teal\",\"drink\":\"kvass\"}}</answer>"}
{"model": "nvidia/switchyard", "task": "E1", "tier": "easy", "rep": 2, "pass": true, "detail": "{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 7.8678e-05, "in_tok": 158, "out_tok": 100, "latency": 2.63, "id": "gen-1791297773-tv3MyNvUAhToGGWz7w50", "text": "<answer>{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}</answer>"}
{"model": "nvidia/switchyard", "task": "E2", "tier": "easy", "rep": 2, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "deepseek/deepseek-v4.1-flash", "provider": "DigitalOcean", "cost": 0.00018828, "in_tok": 178, "out_tok": 217, "latency": 1.71, "id": "gen-1791297775-7NEH7vpklqUcL0H2v2Uo", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "nvidia/switchyard", "task": "H4", "tier": "hard", "rep": 0, "pass": true, "detail": "grid ok", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.070896, "in_tok": 366, "out_tok": 4641, "latency": 38.77, "id": "gen-1791297738-ILHZu3TAtehzuazUYoSz", "text": "<answer>{\"1\":{\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\"2\":{\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\"3\":{\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\"4\":{\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\"5\":{\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "nvidia/switchyard", "task": "H2", "tier": "hard", "rep": 0, "pass": true, "detail": "46/46 hidden tests", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.063765, "in_tok": 300, "out_tok": 4181, "latency": 36.07, "id": "gen-1791297742-0Q77LOVV0dFVFmcKw2Af", "text": "```python\ndef vault(grid: list[str]) -> int:\n    import heapq\n\n    n = len(grid)\n    m = len(grid[0]) if n else 0\n    start = end = None\n    stars = []\n\n    for r in range(n):\n        for c in range(m):\n            ch = grid[r][c]\n            if ch == \"S\":\n                start = (r, c)\n            elif ch == \"E\":\n                end = (r, c)\n            elif ch == \"*\":\n                stars.append((r, c))\n\n    if start is None or end is None:\n        return -1\n\n    jump = {}\n    if len(stars) == 2:\n        a, b = stars\n        jump[a] = b\n        jump[b] = a\n\n    INF = 10**9\n    sr, sc = start\n    dist = {(sr, sc, 0): 0}\n    pq = [(0, sr, sc, 0)]\n\n    while pq:\n        d, r, c, mask = heapq.heappop(pq)\n        if d != dist.get((r, c, mask)):\n            continue\n        if (r, c) == end:\n            return d\n\n        other = jump.get((r, c))\n        if other is not None:\n            nr, nc = other\n            state = (nr, nc, mask)\n            if d < dist.get(state, INF):\n                dist[state] = d\n                heapq.heappush(pq, (d, nr, nc, mask))\n\n        for dr, dc in ((1, 0), (-1, 0), (0, 1), (0, -1)):\n            nr, nc = r + dr, c + dc\n            if nr < 0 or nr >= n or nc < 0 or nc >= m:\n                continue\n            ch = grid[nr][nc]\n            if ch == \"#\":\n                continue\n            if \"A\" <= ch <= \"C\" and not (mask >> (ord(ch) - 65)) & 1:\n                continue\n\n            nmask = mask\n            if \"a\" <= ch <= \"c\":\n                nmask |= 1 << (ord(ch) - 97)\n\n            nd = d + 1 + (ord(ch) - 48 if \"1\" <= ch <= \"9\" else 0)\n            state = (nr, nc, nmask)\n            if nd < dist.get(state, INF):\n                dist[state] = nd\n                heapq.heappush(pq, (nd, nr, nc, nmask))\n\n    return -1\n```"}
{"model": "nvidia/switchyard", "task": "E4", "tier": "easy", "rep": 2, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0002994, "in_tok": 102, "out_tok": 224, "latency": 2.38, "id": "gen-1791297777-rrFfxki0cK88KwwCkDos", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "nvidia/switchyard", "task": "E3", "tier": "easy", "rep": 2, "pass": true, "detail": "520", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0033615, "in_tok": 159, "out_tok": 187, "latency": 3.18, "id": "gen-1791297777-YuLvTF65I9syexEURLi8", "text": "<answer>520</answer>"}
{"model": "nvidia/switchyard", "task": "H2", "tier": "hard", "rep": 1, "pass": true, "detail": "46/46 hidden tests", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.04638, "in_tok": 300, "out_tok": 3022, "latency": 20.38, "id": "gen-1791297760-E44a65BthELeZRkZRs6o", "text": "```python\ndef vault(grid: list[str]) -> int:\n    import heapq\n\n    R, C = len(grid), len(grid[0])\n    start = end = None\n    stars = []\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            if ch == 'S':\n                start = (r, c)\n            elif ch == 'E':\n                end = (r, c)\n            elif ch == '*':\n                stars.append((r, c))\n\n    portal = {}\n    if len(stars) == 2:\n        portal[stars[0]] = stars[1]\n        portal[stars[1]] = stars[0]\n\n    INF = float('inf')\n    dist = {(start[0], start[1], 0): 0}\n    pq = [(0, start[0], start[1], 0)]\n    DIRS = ((1, 0), (-1, 0), (0, 1), (0, -1))\n\n    while pq:\n        d, r, c, keys = heapq.heappop(pq)\n        if d != dist.get((r, c, keys)):\n            continue\n        if (r, c) == end:\n            return d\n\n        dest = portal.get((r, c))\n        if dest is not None:\n            st = (dest[0], dest[1], keys)\n            if d < dist.get(st, INF):\n                dist[st] = d\n                heapq.heappush(pq, (d, dest[0], dest[1], keys))\n\n        for dr, dc in DIRS:\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n            ch = grid[nr][nc]\n            if ch == '#':\n                continue\n            if 'A' <= ch <= 'C' and not (keys >> (ord(ch) - 65)) & 1:\n                continue\n            nk = keys\n            if 'a' <= ch <= 'c':\n                nk = keys | (1 << (ord(ch) - 97))\n            cost = 1 + (ord(ch) - 48 if '1' <= ch <= '9' else 0)\n            nd = d + cost\n            st = (nr, nc, nk)\n            if nd < dist.get(st, INF):\n                dist[st] = nd\n                heapq.heappush(pq, (nd, nr, nc, nk))\n\n    return -1\n```"}
{"model": "nvidia/switchyard", "task": "M1", "tier": "medium", "rep": 2, "pass": true, "detail": "3, 32, 10, 8, 2, 9", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0166955, "in_tok": 343, "out_tok": 1033, "latency": 9.8, "id": "gen-1791297778-rr2PX66P0OcPkOvvpsxk", "text": "<answer>3, 32, 10, 8, 2, 9</answer>"}
{"model": "nvidia/switchyard", "task": "H4", "tier": "hard", "rep": 1, "pass": true, "detail": "grid ok", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.069396, "in_tok": 366, "out_tok": 4541, "latency": 37.95, "id": "gen-1791297759-yKya3xGxJNWyo11taoub", "text": "<answer>{\"1\": {\"name\": \"Fenna\", \"color\": \"coral\", \"drink\": \"chicha\"}, \"2\": {\"name\": \"Dagny\", \"color\": \"teal\", \"drink\": \"kvass\"}, \"3\": {\"name\": \"Bram\", \"color\": \"slate\", \"drink\": \"horchata\"}, \"4\": {\"name\": \"Evander\", \"color\": \"ochre\", \"drink\": \"sbiten\"}, \"5\": {\"name\": \"Gideon\", \"color\": \"olive\", \"drink\": \"mate\"}}</answer>"}
{"model": "nvidia/switchyard", "task": "M4", "tier": "medium", "rep": 2, "pass": true, "detail": "grid ok", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.019253, "in_tok": 298, "out_tok": 1214, "latency": 11.17, "id": "gen-1791297788-GHDyQYD2ctf7yqnTwNPF", "text": "<answer>{\"1\": {\"name\": \"Ilsa\", \"color\": \"slate\", \"drink\": \"mate\"}, \"2\": {\"name\": \"Bram\", \"color\": \"plum\", \"drink\": \"horchata\"}, \"3\": {\"name\": \"Dagny\", \"color\": \"olive\", \"drink\": \"chicha\"}, \"4\": {\"name\": \"Evander\", \"color\": \"teal\", \"drink\": \"kvass\"}}</answer>"}
{"model": "nvidia/switchyard", "task": "M3", "tier": "medium", "rep": 1, "pass": true, "detail": "27 Dross 1130", "routed": "moonshotai/kimi-k3", "provider": "Morph", "cost": 0.046698462, "in_tok": 230, "out_tok": 3619, "latency": 51.52, "id": "gen-1791297753-N6222RIZeuHRzCxpUnIw", "text": "<answer>27 Dross 1130</answer>"}
{"model": "nvidia/switchyard", "task": "M3", "tier": "medium", "rep": 2, "pass": true, "detail": "27 Dross 1130", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0395695, "in_tok": 227, "out_tok": 2585, "latency": 26.74, "id": "gen-1791297781-akKWvVVC9tHp7il5E7I3", "text": "<answer>27 Dross 1130</answer>"}
{"model": "nvidia/switchyard", "task": "M2", "tier": "medium", "rep": 2, "pass": true, "detail": "37/37 hidden tests", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0460935, "in_tok": 201, "out_tok": 3026, "latency": 30.03, "id": "gen-1791297780-GaFQ5NW5RhJL1gf2jKw0", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    stack: list[int] = []\n    out = [0] * n\n\n    for i in range(n - 1, -1, -1):\n        xi = xs[i]\n        while stack and xs[stack[-1]] < xi:\n            stack.pop()\n        d = stack[-1] - i if stack else 0\n        out[i] = xi + d\n        stack.append(i)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "nvidia/switchyard", "task": "H2", "tier": "hard", "rep": 2, "pass": true, "detail": "46/46 hidden tests", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.02565, "in_tok": 300, "out_tok": 1640, "latency": 14.49, "id": "gen-1791297799-BtH5ExguSuF6LJg99HxX", "text": "```python\nimport heapq\n\ndef vault(grid: list[str]) -> int:\n    R, C = len(grid), len(grid[0])\n    start = end = None\n    stars = []\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            if ch == 'S':\n                start = (r, c)\n            elif ch == 'E':\n                end = (r, c)\n            elif ch == '*':\n                stars.append((r, c))\n\n    # Link portals only when there are exactly two.\n    link = {}\n    if len(stars) == 2:\n        link[stars[0]] = stars[1]\n        link[stars[1]] = stars[0]\n\n    INF = float('inf')\n    dist = [[[INF] * 8 for _ in range(C)] for _ in range(R)]\n    sr, sc = start\n    dist[sr][sc][0] = 0\n    pq = [(0, sr, sc, 0)]  # (cost, row, col, keymask)\n\n    while pq:\n        cost, r, c, mask = heapq.heappop(pq)\n        if cost > dist[r][c][mask]:\n            continue\n        if (r, c) == end:\n            return cost\n\n        # Zero-cost portal jump.\n        dst = link.get((r, c))\n        if dst is not None and cost < dist[dst[0]][dst[1]][mask]:\n            dist[dst[0]][dst[1]][mask] = cost\n            heapq.heappush(pq, (cost, dst[0], dst[1], mask))\n\n        # Normal steps.\n        for dr, dc in ((1, 0), (-1, 0), (0, 1), (0, -1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n            ch = grid[nr][nc]\n            if ch == '#':\n                continue\n            ncost, nmask = cost + 1, mask\n            if '1' <= ch <= '9':\n                ncost += ord(ch) - 48\n            elif 'a' <= ch <= 'c':\n                nmask |= 1 << (ord(ch) - 97)\n            elif 'A' <= ch <= 'C':\n                if not (mask >> (ord(ch) - 65)) & 1:\n                    continue\n            if ncost < dist[nr][nc][nmask]:\n                dist[nr][nc][nmask] = ncost\n                heapq.heappush(pq, (ncost, nr, nc, nmask))\n\n    return -1\n```"}
{"model": "unbiased/pareto", "task": "E2", "tier": "easy", "rep": 0, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00195675, "in_tok": 165, "out_tok": 208, "latency": 3.39, "id": "gen-1791297810-XwJlPDfSppO1qULGNOf0", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "unbiased/pareto", "task": "E1", "tier": "easy", "rep": 0, "pass": true, "detail": "{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00072925, "in_tok": 145, "out_tok": 51, "latency": 8.03, "id": "gen-1791297808-zlAbgphOVpAPm0UyzgvV", "text": "<answer>{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}</answer>"}
{"model": "unbiased/pareto", "task": "E3", "tier": "easy", "rep": 0, "pass": true, "detail": "520", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.0039125, "in_tok": 95, "out_tok": 490, "latency": 5.01, "id": "gen-1791297814-PZDGNgrLTog0OFoMTZo4", "text": "To determine how many crates passed inspection, we need to calculate the total number of crates transported during the week and then subtract the number of crates that were rejected.\n\n**Step 1: Calculate the total capacity for a full week.**\nA normal day consists of 4 trips, and each trip carries 35 crates.\nCrates per day = $35 \\text{ crates/trip} \\times 4 \\text{ trips/day} = 140 \\text{ crates/day}$.\n\nFor a full 4-day week without any skips:\nTotal potential crates = $140 \\text{ crates/day} \\times 4 \\text{ days} = 560 \\text{ crates}$.\n\n**Step 2: Account for the skipped trip.**\nOn the last day of the 4-day week, the ferry skips the final trip. This means on the 4th day, it only made 3 trips instead of 4.\nCrates skipped = $1 \\text{ trip} \\times 35 \\text{ crates/trip} = 35 \\text{ crates}$.\n\nSo, the actual number of crates transported that week is:\nTotal transported = Total potential crates - Crates skipped\nTotal transported = $560 - 35 = 525 \\text{ crates}$.\n\nAlternatively, you can calculate it day by day:\n- Days 1, 2, and 3: $3 \\text{ days} \\times 140 \\text{ crates/day} = 420 \\text{ crates}$.\n- Day 4: $3 \\text{ trips} \\times 35 \\text{ crates/trip} = 105 \\text{ crates}$.\n- Total = $420 + 105 = 525 \\text{ crates}$.\n\n**Step 3: Subtract the rejected crates.**\nWe are told that 5 crates were rejected at inspection over the whole week. These crates are part of the transported amount but did not pass.\nCrates passed = Total transported - Rejected crates\nCrates passed = $525 - 5 = 520$.\n\n**Conclusion:**\n520 crates passed inspection.\n\n<answer>520</answer>"}
{"model": "unbiased/pareto", "task": "E4", "tier": "easy", "rep": 0, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00060425, "in_tok": 89, "out_tok": 53, "latency": 8.18, "id": "gen-1791297814-XXMky5hs3F1Uq8H0cRhb", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "nvidia/switchyard", "task": "H3", "tier": "hard", "rep": 0, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.05336334, "in_tok": 298, "out_tok": 3794, "latency": 81.78, "id": "gen-1791297743-Vr0mCIn8ClUx6ANakblq", "text": "<answer>{\"J1\": 0, \"J2\": 8, \"J3\": 0, \"J4\": 5, \"J5\": 6, \"J6\": 1, \"J7\": 7, \"J8\": 11}</answer>"}
{"model": "unbiased/pareto", "task": "M1", "tier": "medium", "rep": 0, "pass": true, "detail": "3, 32, 10, 8, 2, 9", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00235425, "in_tok": 282, "out_tok": 222, "latency": 12.34, "id": "gen-1791297816-EEqtQNEX9IEVzTMbJgOc", "text": "<answer>3, 32, 10, 8, 2, 9</answer>"}
{"model": "unbiased/pareto", "task": "M2", "tier": "medium", "rep": 0, "pass": true, "detail": "37/37 hidden tests", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.0137525, "in_tok": 125, "out_tok": 1792, "latency": 8.43, "id": "gen-1791297822-U3HtzUxtI97MG16lnTV2", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    out = [0] * n\n    stack = []\n\n    for i in range(n - 1, -1, -1):\n        x = xs[i]\n        while stack and xs[stack[-1]] < x:\n            stack.pop()\n\n        if stack:\n            j = stack[-1]\n            out[i] = x + (j - i)\n        else:\n            out[i] = x\n\n        stack.append(i)\n\n    k = sum(xs) % n\n    if k:\n        out = out[k:] + out[:k]\n\n    return out\n```"}
{"model": "nvidia/switchyard", "task": "H3", "tier": "hard", "rep": 1, "pass": false, "detail": "3 jobs at t=5", "routed": "moonshotai/kimi-k3", "provider": "Morph", "cost": 0.071399955, "in_tok": 301, "out_tok": 5538, "latency": 77.96, "id": "gen-1791297763-T99TCjUePROcIvrGaBtD", "text": "<answer>{\"J1\": 0, \"J2\": 8, \"J3\": 0, \"J4\": 5, \"J5\": 6, \"J6\": 1, \"J7\": 5, \"J8\": 10}</answer>"}
{"model": "unbiased/pareto", "task": "M3", "tier": "medium", "rep": 0, "pass": true, "detail": "27 Dross 1130", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.002175, "in_tok": 162, "out_tok": 236, "latency": 25.79, "id": "gen-1791297825-gWdrQLMqwFT3Sgpj0vQz", "text": "<answer>27 Dross 1130</answer>"}
{"model": "unbiased/pareto", "task": "M4", "tier": "medium", "rep": 0, "pass": true, "detail": "grid ok", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00226925, "in_tok": 239, "out_tok": 225, "latency": 23.46, "id": "gen-1791297828-wDPtqqjwtmViLuFdp33r", "text": "<answer>{\"1\":{\"name\":\"Ilsa\",\"color\":\"slate\",\"drink\":\"mate\"},\"2\":{\"name\":\"Bram\",\"color\":\"plum\",\"drink\":\"horchata\"},\"3\":{\"name\":\"Dagny\",\"color\":\"olive\",\"drink\":\"chicha\"},\"4\":{\"name\":\"Evander\",\"color\":\"teal\",\"drink\":\"kvass\"}}</answer>"}
{"model": "unbiased/pareto", "task": "E1", "tier": "easy", "rep": 1, "pass": true, "detail": "{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00059275, "in_tok": 145, "out_tok": 52, "latency": 9.87, "id": "gen-1791297852-qoP6bUs903zVryWHhYvK", "text": "<answer>{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}</answer>"}
{"model": "unbiased/pareto", "task": "E2", "tier": "easy", "rep": 1, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00127275, "in_tok": 165, "out_tok": 136, "latency": 2.43, "id": "gen-1791297861-9b6u3M8nCBqZ6tzdnhS5", "text": "1. \"Box was crushed and the lamp shade is bent.\" -> Physical item arrived damaged -> TARN\n2. \"Please let me export to CSV.\" -> Feature request -> WREN\n3. \"Why is there a $4 fee on my statement?\" -> Being charged money -> GLINT\n4. \"Could you add a dark mode?\" -> Feature request -> WREN\n5. \"The mug came cracked in the box.\" -> Physical item arrived damaged -> TARN\n6. \"2FA code keeps getting rejected.\" -> Login or password trouble -> MOSS\n\n<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "unbiased/pareto", "task": "H2", "tier": "hard", "rep": 0, "pass": true, "detail": "46/46 hidden tests", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00527175, "in_tok": 240, "out_tok": 625, "latency": 25.03, "id": "gen-1791297841-Q4FuI8U3AJ8xBZBExwiC", "text": "```python\ndef vault(grid: list[str]) -> int:\n    from heapq import heappop, heappush\n\n    if not grid or not grid[0]:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    start = None\n    portals = []\n\n    for r in range(rows):\n        for c in range(cols):\n            if grid[r][c] == \"S\":\n                start = (r, c)\n            elif grid[r][c] == \"*\":\n                portals.append((r, c))\n\n    if start is None:\n        return -1\n\n    links = {}\n    if len(portals) == 2:\n        a, b = portals\n        links[a] = b\n        links[b] = a\n\n    sr, sc = start\n    distances = {(sr, sc, 0): 0}\n    heap = [(0, sr, sc, 0)]\n\n    while heap:\n        cost, r, c, keys = heappop(heap)\n        if cost != distances.get((r, c, keys)):\n            continue\n        if grid[r][c] == \"E\":\n            return cost\n\n        if (r, c) in links:\n            nr, nc = links[(r, c)]\n            state = (nr, nc, keys)\n            if cost < distances.get(state, float(\"inf\")):\n                distances[state] = cost\n                heappush(heap, (cost, nr, nc, keys))\n\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < rows and 0 <= nc < cols):\n                continue\n\n            cell = grid[nr][nc]\n            if cell == \"#\":\n                continue\n            if cell in \"ABC\" and not (keys & (1 << (ord(cell) - ord(\"A\")))):\n                continue\n\n            new_keys = keys\n            if cell in \"abc\":\n                new_keys |= 1 << (ord(cell) - ord(\"a\"))\n\n            new_cost = cost + 1\n            if \"1\" <= cell <= \"9\":\n                new_cost += int(cell)\n\n            state = (nr, nc, new_keys)\n            if new_cost < distances.get(state, float(\"inf\")):\n                distances[state] = new_cost\n                heappush(heap, (new_cost, nr, nc, new_keys))\n\n    return -1\n```"}
{"model": "unbiased/pareto", "task": "E3", "tier": "easy", "rep": 1, "pass": true, "detail": "520", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.0011675, "in_tok": 95, "out_tok": 124, "latency": 5.42, "id": "gen-1791297864-5Rz5WpWssQhoo2Dn19tM", "text": "<answer>520</answer>"}
{"model": "unbiased/pareto", "task": "E4", "tier": "easy", "rep": 1, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00069175, "in_tok": 94, "out_tok": 63, "latency": 6.89, "id": "gen-1791297867-VEOO1DCU7ehh7iAOB1un", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "unbiased/pareto", "task": "M1", "tier": "medium", "rep": 1, "pass": true, "detail": "3, 32, 10, 8, 2, 9", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00188175, "in_tok": 282, "out_tok": 159, "latency": 11.53, "id": "gen-1791297869-0de16KDWbzjXr8G9j1CO", "text": "<answer>3, 32, 10, 8, 2, 9</answer>"}
{"model": "nvidia/switchyard", "task": "H1", "tier": "hard", "rep": 0, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "moonshotai/kimi-k3", "provider": "AkashML", "cost": 0.1387241, "in_tok": 397, "out_tok": 9872, "latency": 145.14, "id": "gen-1791297736-pprKZoYCV4kOeiD3JXo8", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "nvidia/switchyard", "task": "H1", "tier": "hard", "rep": 1, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "moonshotai/kimi-k3", "provider": "AkashML", "cost": 0.1267541, "in_tok": 397, "out_tok": 9017, "latency": 132.92, "id": "gen-1791297752-WpNVb7hrvYNHze0ssWzo", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "nvidia/switchyard", "task": "H4", "tier": "hard", "rep": 2, "pass": true, "detail": "grid ok", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.123891, "in_tok": 366, "out_tok": 8174, "latency": 93.53, "id": "gen-1791297797-LuMwYPscmxQkfWnQZEPG", "text": "<answer>{\"1\":{\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\"2\":{\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\"3\":{\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\"4\":{\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\"5\":{\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "unbiased/pareto", "task": "H1", "tier": "hard", "rep": 0, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00491175, "in_tok": 336, "out_tok": 545, "latency": 73.29, "id": "gen-1791297819-F7lntmI7pM5w11LHPOKc", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "unbiased/pareto", "task": "M2", "tier": "medium", "rep": 1, "pass": true, "detail": "37/37 hidden tests", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.0010325, "in_tok": 125, "out_tok": 96, "latency": 13.54, "id": "gen-1791297881-bGZ923N0DhBRNhy5keod", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    if not xs:\n        return []\n\n    out = []\n    for i in range(len(xs)):\n        d = next((j - i for j in range(i + 1, len(xs)) if xs[j] >= xs[i]), 0)\n        out.append(xs[i] + d)\n\n    k = sum(xs) % len(xs)\n    return out[k:] + out[:k]\n```"}
{"model": "unbiased/pareto", "task": "H4", "tier": "hard", "rep": 0, "pass": true, "detail": "grid ok", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00438175, "in_tok": 307, "out_tok": 484, "latency": 68.92, "id": "gen-1791297831-bjXq2oC5M2Jdp0qoXrEp", "text": "<answer>{\"1\":{\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\"2\":{\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\"3\":{\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\"4\":{\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\"5\":{\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "unbiased/pareto", "task": "E1", "tier": "easy", "rep": 2, "pass": true, "detail": "{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00098275, "in_tok": 145, "out_tok": 104, "latency": 4.05, "id": "gen-1791297900-WJwIJxe4P9AN1GELzcQ0", "text": "<answer>{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}</answer>"}
{"model": "unbiased/pareto", "task": "H2", "tier": "hard", "rep": 1, "pass": false, "detail": "no code block", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.0042425, "in_tok": 224, "out_tok": 491, "latency": 13.29, "id": "gen-1791297892-UbRnoqgtanV8Jl3xgovJ", "text": "Drafts 1 and 2 are functionally equivalent and correct (Dijkstra over `(r, c, keymask)` states, 0-cost portal links when exactly two `*`, proper digit/door cost/guard handling). Draft 3 uses plain BFS on a weighted graph (digits add extra cost), which is incorrect, and its portal handling is buggy. Since a draft suffices, SELECT draft 1 (cleanest correct implementation).\n\nSELECT: 1"}
{"model": "unbiased/pareto", "task": "E2", "tier": "easy", "rep": 2, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.001875, "in_tok": 162, "out_tok": 196, "latency": 2.15, "id": "gen-1791297904-5hu2lhw2rOZzICZ4zaIE", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "nvidia/switchyard", "task": "H3", "tier": "hard", "rep": 2, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.177758, "in_tok": 298, "out_tok": 11781, "latency": 105.55, "id": "gen-1791297804-FVEjkfwoip6naWQLlll6", "text": "<answer>{\"J1\": 0, \"J2\": 8, \"J3\": 0, \"J4\": 7, \"J5\": 5, \"J6\": 1, \"J7\": 7, \"J8\": 11}</answer>"}
{"model": "unbiased/pareto", "task": "E3", "tier": "easy", "rep": 2, "pass": true, "detail": "520", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00338, "in_tok": 95, "out_tok": 419, "latency": 5.99, "id": "gen-1791297905-LRbKJzzDmmPjMESDcBGM", "text": "To determine how many crates passed inspection, we need to calculate the total number of crates transported during the week and then subtract the number of crates that were rejected.\n\n**Step 1: Calculate the total number of trips made during the week.**\n\n*   The ferry normally makes 4 trips per day.\n*   The week consists of 4 days.\n*   On the last day (the 4th day), it skips the final trip. This means it only makes 3 trips on that day instead of 4.\n*   Trips on the first 3 days: $3 \\text{ days} \\times 4 \\text{ trips/day} = 12 \\text{ trips}$.\n*   Trips on the 4th day: $3 \\text{ trips}$.\n*   Total trips for the week: $12 + 3 = 15 \\text{ trips}$.\n\n**Step 2: Calculate the total number of crates transported.**\n\n*   Each trip carries 35 crates.\n*   Total crates transported = $\\text{Total trips} \\times \\text{Crates per trip}$\n*   Total crates transported = $15 \\times 35$\n\nLet's do the multiplication:\n$15 \\times 35 = 15 \\times (30 + 5) = (15 \\times 30) + (15 \\times 5) = 450 + 75 = 525$.\n\nSo, 525 crates were transported in total.\n\n**Step 3: Subtract the rejected crates to find the passed crates.**\n\n*   Total crates transported: 525\n*   Crates rejected: 5\n*   Crates passed = $\\text{Total transported} - \\text{Rejected}$\n*   Crates passed = $525 - 5 = 520$\n\n<answer>520</answer>"}
{"model": "unbiased/pareto", "task": "M4", "tier": "medium", "rep": 1, "pass": true, "detail": "grid ok", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00211775, "in_tok": 239, "out_tok": 272, "latency": 26.34, "id": "gen-1791297885-Ezs1dv1dqf17dsDuuK0C", "text": "<answer>{\"1\":{\"name\":\"Ilsa\",\"color\":\"slate\",\"drink\":\"mate\"},\"2\":{\"name\":\"Bram\",\"color\":\"plum\",\"drink\":\"horchata\"},\"3\":{\"name\":\"Dagny\",\"color\":\"olive\",\"drink\":\"chicha\"},\"4\":{\"name\":\"Evander\",\"color\":\"teal\",\"drink\":\"kvass\"}}</answer>"}
{"model": "unbiased/pareto", "task": "E4", "tier": "easy", "rep": 2, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00049675, "in_tok": 94, "out_tok": 61, "latency": 6.79, "id": "gen-1791297906-341DHjjWktHVr0oEjV2n", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "unbiased/pareto", "task": "M3", "tier": "medium", "rep": 1, "pass": true, "detail": "27 Dross 1130", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00197775, "in_tok": 162, "out_tok": 231, "latency": 36.25, "id": "gen-1791297881-cB7v1BvoilXGMWTbZAg4", "text": "<answer>27 Dross 1130</answer>"}
{"model": "unbiased/pareto", "task": "M2", "tier": "medium", "rep": 2, "pass": true, "detail": "37/37 hidden tests", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.0088325, "in_tok": 125, "out_tok": 1136, "latency": 6.41, "id": "gen-1791297911-5thHFR0otuZ7oZDx6j5N", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    out = [0] * n\n    stack = []\n\n    for i in range(n - 1, -1, -1):\n        x = xs[i]\n        while stack and xs[stack[-1]] < x:\n            stack.pop()\n        if stack:\n            out[i] = x + (stack[-1] - i)\n        else:\n            out[i] = x\n        stack.append(i)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "unbiased/pareto", "task": "M1", "tier": "medium", "rep": 2, "pass": true, "detail": "3, 32, 10, 8, 2, 9", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.01112475, "in_tok": 282, "out_tok": 1473, "latency": 9.02, "id": "gen-1791297910-rLrXUGAm9FreENKT1vR6", "text": "To determine the final stack of the GLYPH program, we will execute the instructions step by step, maintaining the stack state. The stack grows to the right (top is the last element).\n\n**Program:** `3 4 8 swap over mul 10 4 11 swap sub 4 mul mod rot 2 9`\n\n**Step-by-Step Execution:**\n\n1.  **`3`**: Push 3.\n    *   Stack: `[3]`\n\n2.  **`4`**: Push 4.\n    *   Stack: `[3, 4]`\n\n3.  **`8`**: Push 8.\n    *   Stack: `[3, 4, 8]`\n\n4.  **`swap`**: Swap top two elements (8 and 4).\n    *   Stack: `[3, 8, 4]`\n\n5.  **`over`**: Copy second-from-top to top. Top is 4, second-from-top is 8. Push 8.\n    *   Stack: `[3, 8, 4, 8]`\n\n6.  **`mul`**: Pop top two (8, 4), push $4 \\times 8 = 32$. Note: `mul` pops `b` then `a`? The definition says \"pop b then a, push a*b\". Usually, in stack operations, the first pop is the top element ($b$), and the second pop is the new top ($a$). So we compute $a \\times b$. Here $a=4, b=8$. $4 \\times 8 = 32$.\n    *   Stack: `[3, 8, 32]`\n\n7.  **`10`**: Push 10.\n    *   Stack: `[3, 8, 32, 10]`\n\n8.  **`4`**: Push 4.\n    *   Stack: `[3, 8, 32, 10, 4]`\n\n9.  **`11`**: Push 11.\n    *   Stack: `[3, 8, 32, 10, 4, 11]`\n\n10. **`swap`**: Swap top two (11 and 4).\n    *   Stack: `[3, 8, 32, 10, 11, 4]`\n\n11. **`sub`**: Pop top two ($b=4, a=11$), push $a - b = 11 - 4 = 7$.\n    *   Stack: `[3, 8, 32, 10, 7]`\n\n12. **`4`**: Push 4.\n    *   Stack: `[3, 8, 32, 10, 7, 4]`\n\n13. **`mul`**: Pop top two ($b=4, a=7$), push $a \\times b = 7 \\times 4 = 28$.\n    *   Stack: `[3, 8, 32, 10, 28]`\n\n14. **`mod`**: Pop top two ($b=28, a=10$). Push $a \\pmod b = 10 \\pmod{28}$.\n    *   $10 \\div 28 = 0$ with remainder $10$.\n    *   Stack: `[3, 8, 32, 10]`\n\n15. **`rot`**: Rotate top three. Definition: `(a b c -> b c a)`. This means the third-from-top moves to the top.\n    *   Current top three: $32$ (bottom of triplet), $10$ (middle), $10$ (top)? Wait, let's look at the stack `[3, 8, 32, 10]`.\n    *   Top is 10. Second is 32. Third is 8.\n    *   Let's re-read carefully: `mod` result was 10. Stack before rot: `[3, 8, 32, 10]`.\n    *   Top ($c$) = 10. Second ($b$) = 32. Third ($a$) = 8.\n    *   `rot` transforms $(a, b, c)$ to $(b, c, a)$.\n    *   So, 8 moves to top. The new top three become 32, 10, 8.\n    *   Stack: `[3, 8, 10, 32]`? No.\n    *   Let's trace standard `rot` (often defined as removing top 3 and pushing b, c, a).\n    *   Stack: `[3, 8, 32, 10]`.\n    *   Top 3 are `8, 32, 10` (where 8 is lowest of the three).\n    *   $a=8, b=32, c=10$.\n    *   Result order pushed: $b, c, a \\rightarrow 32, 10, 8$.\n    *   So the stack becomes `[3, 32, 10, 8]`.\n\n    Let's double check the `rot` definition provided: `(a b c -> b c a)`.\n    In stack notation, if the stack is `[..., a, b, c]`, then `rot` produces `[..., b, c, a]`.\n    Current stack: `[3, 8, 32, 10]`.\n    Here `a=8`, `b=32`, `c=10`.\n    New stack section: `32, 10, 8`.\n    Full stack: `[3, 32, 10, 8]`.\n\n16. **`2`**: Push 2.\n    *   Stack: `[3, 32, 10, 8, 2]`\n\n17. **`9`**: Push 9.\n    *   Stack: `[3, 32, 10, 8, 2, 9]`\n\nThere are no more instructions.\n\nThe final stack, from bottom to top, is: 3, 32, 10, 8, 2, 9.\n\n<answer>3, 32, 10, 8, 2, 9</answer>"}
{"model": "unbiased/pareto", "task": "H3", "tier": "hard", "rep": 0, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00287675, "in_tok": 239, "out_tok": 306, "latency": 70.88, "id": "gen-1791297851-b1e0Xd3yatl35chTRDsO", "text": "<answer>{\"J1\":0,\"J2\":8,\"J3\":0,\"J4\":5,\"J5\":6,\"J6\":1,\"J7\":7,\"J8\":11}</answer>"}
{"model": "unbiased/pareto", "task": "M4", "tier": "medium", "rep": 2, "pass": true, "detail": "grid ok", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00244175, "in_tok": 239, "out_tok": 248, "latency": 19.64, "id": "gen-1791297917-uWjFPm4FulIuS0Mu2JKm", "text": "<answer>{\"1\":{\"name\":\"Ilsa\",\"color\":\"slate\",\"drink\":\"mate\"},\"2\":{\"name\":\"Bram\",\"color\":\"plum\",\"drink\":\"horchata\"},\"3\":{\"name\":\"Dagny\",\"color\":\"olive\",\"drink\":\"chicha\"},\"4\":{\"name\":\"Evander\",\"color\":\"teal\",\"drink\":\"kvass\"}}</answer>"}
{"model": "openrouter/auto", "task": "E1", "tier": "easy", "rep": 0, "pass": true, "detail": "{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}", "routed": "deepseek/deepseek-v4-flash-0731", "provider": "Baidu", "cost": 0.00032516, "in_tok": 211, "out_tok": 176, "latency": 3.96, "id": "gen-1791297937-cHkVtn4mFjajQGX8Canz", "text": "<answer>{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}</answer>"}
{"model": "unbiased/pareto", "task": "H2", "tier": "hard", "rep": 2, "pass": true, "detail": "46/46 hidden tests", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00487725, "in_tok": 237, "out_tok": 631, "latency": 27.03, "id": "gen-1791297919-9SvgOBSGAiAPDv5iAd1V", "text": "```python\ndef vault(grid: list[str]) -> int:\n    from heapq import heappop, heappush\n\n    if not grid or not grid[0]:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    start = end = None\n    portals = []\n\n    for r, row in enumerate(grid):\n        for c, cell in enumerate(row):\n            if cell == \"S\":\n                start = (r, c)\n            elif cell == \"E\":\n                end = (r, c)\n            elif cell == \"*\":\n                portals.append((r, c))\n\n    if start is None or end is None:\n        return -1\n\n    links = {}\n    if len(portals) == 2:\n        a, b = portals\n        links[a] = b\n        links[b] = a\n\n    initial = (start[0], start[1], 0)\n    distances = {initial: 0}\n    heap = [(0, *initial)]\n\n    while heap:\n        cost, r, c, keys = heappop(heap)\n        if cost != distances.get((r, c, keys)):\n            continue\n        if (r, c) == end:\n            return cost\n\n        if (r, c) in links:\n            nr, nc = links[(r, c)]\n            state = (nr, nc, keys)\n            if cost < distances.get(state, float(\"inf\")):\n                distances[state] = cost\n                heappush(heap, (cost, *state))\n\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < rows and 0 <= nc < cols):\n                continue\n\n            cell = grid[nr][nc]\n            if cell == \"#\":\n                continue\n            if cell in \"ABC\" and not (keys & (1 << (ord(cell) - ord(\"A\")))):\n                continue\n\n            next_keys = keys\n            if cell in \"abc\":\n                next_keys |= 1 << (ord(cell) - ord(\"a\"))\n\n            next_cost = cost + 1\n            if cell in \"123456789\":\n                next_cost += int(cell)\n\n            state = (nr, nc, next_keys)\n            if next_cost < distances.get(state, float(\"inf\")):\n                distances[state] = next_cost\n                heappush(heap, (next_cost, *state))\n\n    return -1\n```"}
{"model": "openrouter/auto", "task": "E2", "tier": "easy", "rep": 0, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "deepseek/deepseek-v4-flash-0731", "provider": "Baidu", "cost": 0.00039204, "in_tok": 231, "out_tok": 220, "latency": 6.58, "id": "gen-1791297941-5FQWtHuKRuULW7O765QR", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "openrouter/auto", "task": "E3", "tier": "easy", "rep": 0, "pass": true, "detail": "520", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.000174, "in_tok": 108, "out_tok": 118, "latency": 1.77, "id": "gen-1791297946-YA10sfbn0qDe9VGXINtY", "text": "<answer>520</answer>"}
{"model": "unbiased/pareto", "task": "M3", "tier": "medium", "rep": 2, "pass": true, "detail": "27 Dross 1130", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00189525, "in_tok": 162, "out_tok": 220, "latency": 38.04, "id": "gen-1791297913-r5gq8Y7SHC4kAiyh1Au6", "text": "<answer>27 Dross 1130</answer>"}
{"model": "openrouter/auto", "task": "M1", "tier": "medium", "rep": 0, "pass": true, "detail": "3,32,10,8,2,9", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0013833, "in_tok": 291, "out_tok": 1080, "latency": 4.56, "id": "gen-1791297948-1YZ6o6h85B1Yk7wSfJZl", "text": "<answer>3,32,10,8,2,9</answer>"}
{"model": "openrouter/auto", "task": "E4", "tier": "easy", "rep": 0, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "openai/gpt-6-luna", "provider": "Azure", "cost": 6.35e-05, "in_tok": 75, "out_tok": 112, "latency": 5.98, "id": "gen-1791297948-1d5IuyiX2BIhCUGkE2iB", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "unbiased/pareto", "task": "H1", "tier": "hard", "rep": 1, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00459675, "in_tok": 336, "out_tok": 599, "latency": 82.94, "id": "gen-1791297873-6AriOqkxVS0Yrh0CHbDi", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "nvidia/switchyard", "task": "H1", "tier": "hard", "rep": 2, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "moonshotai/kimi-k3", "provider": "Relace", "cost": 0.13104054, "in_tok": 397, "out_tok": 10055, "latency": 178.71, "id": "gen-1791297780-ynbB7NnRejAfibKjZd0V", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "unbiased/pareto", "task": "H4", "tier": "hard", "rep": 1, "pass": true, "detail": "grid ok", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00518125, "in_tok": 307, "out_tok": 677, "latency": 70.19, "id": "gen-1791297891-YCA5PiyEYQ1FAQ8mMEmo", "text": "<answer>{\"1\":{\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\"2\":{\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\"3\":{\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\"4\":{\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\"5\":{\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "openrouter/auto", "task": "M4", "tier": "medium", "rep": 0, "pass": true, "detail": "grid ok", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0016353, "in_tok": 247, "out_tok": 1301, "latency": 6.35, "id": "gen-1791297956-U3GZdk4AFLU1QsmKgn7f", "text": "<answer>{\"1\":{\"name\":\"Ilsa\",\"color\":\"slate\",\"drink\":\"mate\"},\"2\":{\"name\":\"Bram\",\"color\":\"plum\",\"drink\":\"horchata\"},\"3\":{\"name\":\"Dagny\",\"color\":\"olive\",\"drink\":\"chicha\"},\"4\":{\"name\":\"Evander\",\"color\":\"teal\",\"drink\":\"kvass\"}}</answer>"}
{"model": "openrouter/auto", "task": "M2", "tier": "medium", "rep": 0, "pass": true, "detail": "37/37 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0029289, "in_tok": 151, "out_tok": 2403, "latency": 10.17, "id": "gen-1791297953-md8Y7ZErmIKayNyy5mgI", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    out = [0] * n\n    stack = []\n\n    for i in range(n - 1, -1, -1):\n        while stack and xs[stack[-1]] < xs[i]:\n            stack.pop()\n\n        if stack:\n            d = stack[-1] - i\n        else:\n            d = 0\n\n        out[i] = xs[i] + d\n        stack.append(i)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "openrouter/auto", "task": "E1", "tier": "easy", "rep": 1, "pass": true, "detail": "{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}", "routed": "deepseek/deepseek-v4-flash-0731", "provider": "BaseTen", "cost": 5.382e-05, "in_tok": 212, "out_tok": 101, "latency": 1.19, "id": "gen-1791297963-N52UWB5lRL95eeaYtFJX", "text": "<answer>{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}</answer>"}
{"model": "unbiased/pareto", "task": "H3", "tier": "hard", "rep": 1, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00263525, "in_tok": 239, "out_tok": 341, "latency": 70.82, "id": "gen-1791297895-LoeNgl9uEAF34mHlobpo", "text": "<answer>{\"J1\":0,\"J2\":6,\"J3\":0,\"J4\":5,\"J5\":8,\"J6\":1,\"J7\":7,\"J8\":11}</answer>"}
{"model": "openrouter/auto", "task": "E3", "tier": "easy", "rep": 1, "pass": true, "detail": "520", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0001632, "in_tok": 108, "out_tok": 109, "latency": 0.63, "id": "gen-1791297965-Ns4dwddavWXEPrCJrnGG", "text": "<answer>520</answer>"}
{"model": "openrouter/auto", "task": "M3", "tier": "medium", "rep": 0, "pass": true, "detail": "27 Dross 1130", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0039525, "in_tok": 175, "out_tok": 3250, "latency": 13.5, "id": "gen-1791297954-Ml1mfWVGvrxQBEFOXPl5", "text": "<answer>27 Dross 1130</answer>"}
{"model": "unbiased/pareto", "task": "H4", "tier": "hard", "rep": 2, "pass": true, "detail": "grid ok", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00591925, "in_tok": 307, "out_tok": 689, "latency": 49.17, "id": "gen-1791297918-9WGpJ1ppyROFj39yDL92", "text": "<answer>{\"1\":{\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\"2\":{\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\"3\":{\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\"4\":{\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\"5\":{\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "openrouter/auto", "task": "E2", "tier": "easy", "rep": 1, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "deepseek/deepseek-v4-flash-0731", "provider": "Baidu", "cost": 0.0003828, "in_tok": 231, "out_tok": 213, "latency": 3.38, "id": "gen-1791297964-qAIT5l0O5TV549JguBuV", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "openrouter/auto", "task": "E4", "tier": "easy", "rep": 1, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "openai/gpt-6-luna", "provider": "Azure", "cost": 6.1e-05, "in_tok": 75, "out_tok": 107, "latency": 3.81, "id": "gen-1791297966-kSjNW1GgfigOhyoIZlOS", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "openrouter/auto", "task": "M1", "tier": "medium", "rep": 1, "pass": true, "detail": "3,32,10,8,2,9", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.001293636, "in_tok": 291, "out_tok": 1068, "latency": 3.76, "id": "gen-1791297967-McmOWfVEjo7CAqBNFlgu", "text": "<answer>3,32,10,8,2,9</answer>"}
{"model": "openrouter/auto", "task": "H2", "tier": "hard", "rep": 0, "pass": true, "detail": "46/46 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0033354, "in_tok": 250, "out_tok": 2717, "latency": 11.67, "id": "gen-1791297961-XruhA1T2UeB9EVj4vZXt", "text": "```python\ndef vault(grid: list[str]) -> int:\n    import heapq\n\n    if not grid or not grid[0]:\n        return -1\n\n    R, C = len(grid), len(grid[0])\n\n    start = None\n    end = None\n    portals = []\n\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            if ch == 'S':\n                start = (r, c)\n            elif ch == 'E':\n                end = (r, c)\n            elif ch == '*':\n                portals.append((r, c))\n\n    if start is None or end is None:\n        return -1\n\n    portal_map = {}\n    if len(portals) == 2:\n        a, b = portals\n        portal_map[a] = b\n        portal_map[b] = a\n\n    INF = 10**9\n    dist = [[[INF] * 8 for _ in range(C)] for _ in range(R)]\n\n    sr, sc = start\n    dist[sr][sc][0] = 0\n    heap = [(0, sr, sc, 0)]\n\n    dirs = [(1, 0), (-1, 0), (0, 1), (0, -1)]\n\n    while heap:\n        d, r, c, mask = heapq.heappop(heap)\n\n        if d != dist[r][c][mask]:\n            continue\n\n        if (r, c) == end:\n            return d\n\n        if (r, c) in portal_map:\n            nr, nc = portal_map[(r, c)]\n            if d < dist[nr][nc][mask]:\n                dist[nr][nc][mask] = d\n                heapq.heappush(heap, (d, nr, nc, mask))\n\n        for dr, dc in dirs:\n            nr, nc = r + dr, c + dc\n\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n\n            ch = grid[nr][nc]\n\n            if ch == '#':\n                continue\n\n            nmask = mask\n\n            if 'A' <= ch <= 'C':\n                bit = ord(ch) - ord('A')\n                if not (mask & (1 << bit)):\n                    continue\n            elif 'a' <= ch <= 'c':\n                bit = ord(ch) - ord('a')\n                nmask = mask | (1 << bit)\n\n            cost = 1\n            if '1' <= ch <= '9':\n                cost += ord(ch) - ord('0')\n\n            nd = d + cost\n            if nd < dist[nr][nc][nmask]:\n                dist[nr][nc][nmask] = nd\n                heapq.heappush(heap, (nd, nr, nc, nmask))\n\n    return -1\n```"}
{"model": "openrouter/auto", "task": "M4", "tier": "medium", "rep": 1, "pass": true, "detail": "grid ok", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0017229, "in_tok": 247, "out_tok": 1374, "latency": 4.99, "id": "gen-1791297971-jHZn59rigSJOwa0ntsT9", "text": "<answer>{\"1\":{\"name\":\"Ilsa\",\"color\":\"slate\",\"drink\":\"mate\"},\"2\":{\"name\":\"Bram\",\"color\":\"plum\",\"drink\":\"horchata\"},\"3\":{\"name\":\"Dagny\",\"color\":\"olive\",\"drink\":\"chicha\"},\"4\":{\"name\":\"Evander\",\"color\":\"teal\",\"drink\":\"kvass\"}}</answer>"}
{"model": "openrouter/auto", "task": "M2", "tier": "medium", "rep": 1, "pass": true, "detail": "37/37 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0027561, "in_tok": 151, "out_tok": 2259, "latency": 8.34, "id": "gen-1791297967-cYmTo5Y2gI8iRYSuuzEW", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    out = [0] * n\n    stack = []\n\n    for i in range(n - 1, -1, -1):\n        x = xs[i]\n        while stack and xs[stack[-1]] < x:\n            stack.pop()\n\n        if stack:\n            out[i] = x + (stack[-1] - i)\n        else:\n            out[i] = x\n\n        stack.append(i)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "openrouter/auto", "task": "H3", "tier": "hard", "rep": 0, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0045444, "in_tok": 244, "out_tok": 3726, "latency": 16.08, "id": "gen-1791297963-EaPbULrw27YAme4iySjl", "text": "<answer>{\"J1\":0,\"J2\":8,\"J3\":0,\"J4\":5,\"J5\":6,\"J6\":1,\"J7\":7,\"J8\":11}</answer>"}
{"model": "openrouter/auto", "task": "E1", "tier": "easy", "rep": 2, "pass": true, "detail": "{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}", "routed": "deepseek/deepseek-v4-flash-0731", "provider": "Baidu", "cost": 0.00022088, "in_tok": 211, "out_tok": 97, "latency": 1.38, "id": "gen-1791297979-kMeQZzIWDe4so2ktpTho", "text": "<answer>{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}</answer>"}
{"model": "unbiased/pareto", "task": "H3", "tier": "hard", "rep": 2, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00286925, "in_tok": 239, "out_tok": 305, "latency": 59.82, "id": "gen-1791297921-Fv6Wt78zt6hSU7t0HvE7", "text": "<answer>{\"J1\":0,\"J2\":6,\"J3\":0,\"J4\":5,\"J5\":8,\"J6\":1,\"J7\":7,\"J8\":11}</answer>"}
{"model": "openrouter/auto", "task": "E3", "tier": "easy", "rep": 2, "pass": true, "detail": "520", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0001596, "in_tok": 108, "out_tok": 106, "latency": 0.64, "id": "gen-1791297981-pSha79oQykujEynUi9Yu", "text": "<answer>520</answer>"}
{"model": "openrouter/auto", "task": "H4", "tier": "hard", "rep": 0, "pass": true, "detail": "grid ok", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0072399, "in_tok": 313, "out_tok": 5955, "latency": 23.98, "id": "gen-1791297958-JAMUb8YQgoSRwQDhzWzz", "text": "<answer>{\"1\":{\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\"2\":{\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\"3\":{\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\"4\":{\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\"5\":{\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "openrouter/auto", "task": "M3", "tier": "medium", "rep": 1, "pass": true, "detail": "27 Dross 1130", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0048345, "in_tok": 175, "out_tok": 3985, "latency": 15.27, "id": "gen-1791297970-UElgI6DN7DZx6Yvw57tb", "text": "<answer>27 Dross 1130</answer>"}
{"model": "openrouter/auto", "task": "E4", "tier": "easy", "rep": 2, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "openai/gpt-6-luna", "provider": "Azure", "cost": 7.55e-05, "in_tok": 75, "out_tok": 136, "latency": 3.7, "id": "gen-1791297982-mSt9IIbmd2QnKK9HoWW7", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "openrouter/auto", "task": "M1", "tier": "medium", "rep": 2, "pass": true, "detail": "3,32,10,8,2,9", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.001058436, "in_tok": 291, "out_tok": 872, "latency": 4.4, "id": "gen-1791297984-0mPEIoCQUfTRXmSZH7hZ", "text": "<answer>3,32,10,8,2,9</answer>"}
{"model": "openrouter/auto", "task": "H2", "tier": "hard", "rep": 1, "pass": true, "detail": "46/46 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.004161, "in_tok": 250, "out_tok": 3405, "latency": 12.08, "id": "gen-1791297976-LZRRRzEtJT1ALxfMc3mG", "text": "```python\ndef vault(grid: list[str]) -> int:\n    import heapq\n\n    if not grid or not grid[0]:\n        return -1\n\n    R, C = len(grid), len(grid[0])\n    start = end = None\n    portals = []\n\n    for r in range(R):\n        for c, ch in enumerate(grid[r]):\n            if ch == 'S':\n                start = (r, c)\n            elif ch == 'E':\n                end = (r, c)\n            elif ch == '*':\n                portals.append((r, c))\n\n    if start is None or end is None:\n        return -1\n\n    INF = 10**9\n    dist = [[[INF] * 8 for _ in range(C)] for _ in range(R)]\n\n    sr, sc = start\n    dist[sr][sc][0] = 0\n    pq = [(0, sr, sc, 0)]\n    dirs = [(1, 0), (-1, 0), (0, 1), (0, -1)]\n    has_portals = len(portals) == 2\n\n    while pq:\n        d, r, c, mask = heapq.heappop(pq)\n\n        if d != dist[r][c][mask]:\n            continue\n\n        if (r, c) == end:\n            return d\n\n        for dr, dc in dirs:\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n\n            ch = grid[nr][nc]\n            if ch == '#':\n                continue\n\n            nmask = mask\n            if ch in 'abc':\n                nmask |= 1 << (ord(ch) - ord('a'))\n            elif ch in 'ABC':\n                bit = 1 << (ord(ch) - ord('A'))\n                if not (mask & bit):\n                    continue\n\n            cost = 1\n            if '1' <= ch <= '9':\n                cost += ord(ch) - ord('0')\n\n            nd = d + cost\n            if nd < dist[nr][nc][nmask]:\n                dist[nr][nc][nmask] = nd\n                heapq.heappush(pq, (nd, nr, nc, nmask))\n\n        if has_portals:\n            if (r, c) == portals[0]:\n                nr, nc = portals[1]\n            elif (r, c) == portals[1]:\n                nr, nc = portals[0]\n            else:\n                nr = nc = -1\n\n            if nr != -1:\n                if d < dist[nr][nc][mask]:\n                    dist[nr][nc][mask] = d\n                    heapq.heappush(pq, (d, nr, nc, mask))\n\n    return -1\n```"}
{"model": "openrouter/auto", "task": "E2", "tier": "easy", "rep": 2, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "deepseek/deepseek-v4-flash-0731", "provider": "BaseTen", "cost": 7.722e-05, "in_tok": 232, "out_tok": 181, "latency": 9.42, "id": "gen-1791297980-EE47W95nsdM8Z3ziUX4W", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "openrouter/auto", "task": "M2", "tier": "medium", "rep": 2, "pass": true, "detail": "37/37 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0018753, "in_tok": 151, "out_tok": 1525, "latency": 5.17, "id": "gen-1791297986-6Gy6TXQTU5HwhmSkFHKB", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    out = [0] * n\n    stack = []\n\n    for i in range(n - 1, -1, -1):\n        while stack and xs[stack[-1]] < xs[i]:\n            stack.pop()\n\n        if stack:\n            out[i] = xs[i] + (stack[-1] - i)\n        else:\n            out[i] = xs[i]\n\n        stack.append(i)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "openrouter/auto", "task": "H3", "tier": "hard", "rep": 1, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0051024, "in_tok": 244, "out_tok": 4191, "latency": 15.54, "id": "gen-1791297976-YKZhcgS2h7lwglVDRmEp", "text": "<answer>{\"J1\":0,\"J2\":8,\"J3\":0,\"J4\":5,\"J5\":6,\"J6\":1,\"J7\":7,\"J8\":11}</answer>"}
{"model": "openrouter/auto", "task": "M4", "tier": "medium", "rep": 2, "pass": true, "detail": "grid ok", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0015837, "in_tok": 247, "out_tok": 1258, "latency": 4.86, "id": "gen-1791297988-BwXZoPxFJF5urLcCNn4C", "text": "<answer>{\"1\":{\"name\":\"Ilsa\",\"color\":\"slate\",\"drink\":\"mate\"},\"2\":{\"name\":\"Bram\",\"color\":\"plum\",\"drink\":\"horchata\"},\"3\":{\"name\":\"Dagny\",\"color\":\"olive\",\"drink\":\"chicha\"},\"4\":{\"name\":\"Evander\",\"color\":\"teal\",\"drink\":\"kvass\"}}</answer>"}
{"model": "openrouter/auto", "task": "H1", "tier": "hard", "rep": 0, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0164175, "in_tok": 345, "out_tok": 13595, "latency": 43.35, "id": "gen-1791297951-6uW12ruoGre2U6J0ai4n", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "E1", "tier": "easy", "rep": 0, "pass": true, "detail": "{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Modal", "cost": 0.0001686, "in_tok": 158, "out_tok": 101, "latency": 1.24, "id": "gen-1791297993-T4KggVBcNPb7JFFlCypl", "text": "<answer>{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}</answer>"}
{"model": "unbiased/pareto", "task": "H1", "tier": "hard", "rep": 2, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00533175, "in_tok": 336, "out_tok": 601, "latency": 83.08, "id": "gen-1791297911-PtowrFGuJsi5FSa4tR25", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "E2", "tier": "easy", "rep": 0, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 0.000116466, "in_tok": 178, "out_tok": 162, "latency": 2.07, "id": "gen-1791297994-wkZNruz2buSzYfUk0u3C", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "E3", "tier": "easy", "rep": 0, "pass": true, "detail": "520", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 9.7572e-05, "in_tok": 108, "out_tok": 146, "latency": 1.9, "id": "gen-1791297994-vO5y2Kv3sj1Mob294d4m", "text": "<answer>520</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "E4", "tier": "easy", "rep": 0, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 0.000118722, "in_tok": 102, "out_tok": 185, "latency": 2.01, "id": "gen-1791297994-HgYQlM7nAFTLMRo3ke4E", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "openrouter/auto", "task": "H1", "tier": "hard", "rep": 1, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.012345036, "in_tok": 345, "out_tok": 10264, "latency": 31.37, "id": "gen-1791297967-2eo0M9XIxUZoo5ENemfB", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "M1", "tier": "medium", "rep": 0, "pass": true, "detail": "3,32,10,8,2,9", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.001862436, "in_tok": 291, "out_tok": 1542, "latency": 4.12, "id": "gen-1791297996-CVi2UuWkxAtiiJqUywpi", "text": "<answer>3,32,10,8,2,9</answer>"}
{"model": "openrouter/auto", "task": "H4", "tier": "hard", "rep": 1, "pass": true, "detail": "grid ok", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.009917436, "in_tok": 313, "out_tok": 8249, "latency": 27.67, "id": "gen-1791297973-fTw7WSplLybV2Wft4rCR", "text": "<answer>{\"1\":{\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\"2\":{\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\"3\":{\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\"4\":{\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\"5\":{\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "openrouter/auto", "task": "M3", "tier": "medium", "rep": 2, "pass": true, "detail": "27 Dross 1130", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0051129, "in_tok": 175, "out_tok": 4217, "latency": 15.11, "id": "gen-1791297987-7pkYBC9admDi71fjqoUP", "text": "<answer>27 Dross 1130</answer>"}
{"model": "openrouter/auto", "task": "H2", "tier": "hard", "rep": 2, "pass": true, "detail": "46/46 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Modal", "cost": 0.0042366, "in_tok": 250, "out_tok": 3468, "latency": 13.98, "id": "gen-1791297991-oPxYZIBEO0DKjrm95jl3", "text": "```python\ndef vault(grid: list[str]) -> int:\n    import heapq\n\n    if not grid or not grid[0]:\n        return -1\n\n    R, C = len(grid), len(grid[0])\n    sr = sc = er = ec = -1\n    portals = []\n\n    for r in range(R):\n        for c, ch in enumerate(grid[r]):\n            if ch == 'S':\n                sr, sc = r, c\n            elif ch == 'E':\n                er, ec = r, c\n            elif ch == '*':\n                portals.append((r, c))\n\n    if sr == -1 or er == -1:\n        return -1\n\n    if sr == er and sc == ec:\n        return 0\n\n    key_bit = {'a': 0, 'b': 1, 'c': 2}\n    door_bit = {'A': 1, 'B': 2, 'C': 4}\n\n    portal_other = {}\n    if len(portals) == 2:\n        a, b = portals\n        portal_other[a] = b\n        portal_other[b] = a\n\n    def enter_cost(ch: str) -> int:\n        if '1' <= ch <= '9':\n            return 1 + int(ch)\n        return 1\n\n    INF = 10**18\n    dist = {}\n    start = (sr, sc, 0)\n    dist[start] = 0\n    pq = [(0, sr, sc, 0)]\n\n    dirs = [(1, 0), (-1, 0), (0, 1), (0, -1)]\n\n    while pq:\n        d, r, c, mask = heapq.heappop(pq)\n\n        if dist.get((r, c, mask)) != d:\n            continue\n\n        if r == er and c == ec:\n            return d\n\n        if (r, c) in portal_other:\n            nr, nc = portal_other[(r, c)]\n            state = (nr, nc, mask)\n            if d < dist.get(state, INF):\n                dist[state] = d\n                heapq.heappush(pq, (d, nr, nc, mask))\n\n        for dr, dc in dirs:\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n\n            ch = grid[nr][nc]\n\n            if ch == '#':\n                continue\n\n            if ch in door_bit and not (mask & door_bit[ch]):\n                continue\n\n            nmask = mask\n            if ch in key_bit:\n                nmask |= 1 << key_bit[ch]\n\n            nd = d + enter_cost(ch)\n            state = (nr, nc, nmask)\n\n            if nd < dist.get(state, INF):\n                dist[state] = nd\n                heapq.heappush(pq, (nd, nr, nc, nmask))\n\n    return -1\n```"}
{"model": "openrouter/auto", "task": "H4", "tier": "hard", "rep": 2, "pass": true, "detail": "grid ok", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.006017436, "in_tok": 313, "out_tok": 4999, "latency": 15.78, "id": "gen-1791297990-qWkzUuZoVJ6ANkq1vfCP", "text": "<answer>{\"1\":{\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\"2\":{\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\"3\":{\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\"4\":{\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\"5\":{\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "openrouter/auto", "task": "H3", "tier": "hard", "rep": 2, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Modal", "cost": 0.0053736, "in_tok": 244, "out_tok": 4417, "latency": 13.9, "id": "gen-1791297991-aRjfOj8UlKpENRH91gNb", "text": "<answer>{\"J1\": 0, \"J2\": 8, \"J3\": 4, \"J4\": 5, \"J5\": 6, \"J6\": 5, \"J7\": 0, \"J8\": 11}</answer>"}
{"model": "google/gemini-3.8-flash", "task": "E2", "tier": "easy", "rep": 0, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.000936, "in_tok": 158, "out_tok": 218, "latency": 2.68, "id": "gen-1791298005-Z5aFfgZNOOzpXF0mhIyq", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "google/gemini-3.8-flash", "task": "E1", "tier": "easy", "rep": 0, "pass": true, "detail": "{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.00132675, "in_tok": 139, "out_tok": 326, "latency": 3.33, "id": "gen-1791298005-5lebEvin5E8Gx8736vuD", "text": "<answer>{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "M4", "tier": "medium", "rep": 0, "pass": true, "detail": "grid ok", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 0.001148727, "in_tok": 247, "out_tok": 1975, "latency": 9.6, "id": "gen-1791298000-7lVi3EW0mHOw5Gxwo3q3", "text": "<answer>{\"1\":{\"name\":\"Ilsa\",\"color\":\"slate\",\"drink\":\"mate\"},\"2\":{\"name\":\"Bram\",\"color\":\"plum\",\"drink\":\"horchata\"},\"3\":{\"name\":\"Dagny\",\"color\":\"olive\",\"drink\":\"chicha\"},\"4\":{\"name\":\"Evander\",\"color\":\"teal\",\"drink\":\"kvass\"}}</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "M2", "tier": "medium", "rep": 0, "pass": true, "detail": "37/37 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "InferenceNet", "cost": 0.00059197, "in_tok": 151, "out_tok": 1938, "latency": 13.88, "id": "gen-1791297996-ir1mQmSovfdg5N6TSAgr", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    out = [0] * n\n    stack = []\n\n    for i in range(n - 1, -1, -1):\n        x = xs[i]\n        while stack and xs[stack[-1]] < x:\n            stack.pop()\n\n        if stack:\n            out[i] = x + (stack[-1] - i)\n        else:\n            out[i] = x\n\n        stack.append(i)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "google/gemini-3.8-flash", "task": "E3", "tier": "easy", "rep": 0, "pass": true, "detail": "520", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.00071025, "in_tok": 77, "out_tok": 174, "latency": 2.79, "id": "gen-1791298008-nOelRfFQPYnpLeh86t05", "text": "<answer>520</answer>"}
{"model": "google/gemini-3.8-flash", "task": "E4", "tier": "easy", "rep": 0, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.00134775, "in_tok": 67, "out_tok": 346, "latency": 3.9, "id": "gen-1791298009-OljyJJ3NSvsqSiImuLQY", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "openrouter/auto", "task": "H1", "tier": "hard", "rep": 2, "pass": true, "detail": "3,87846,8,4,-586,-571", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.014344236, "in_tok": 345, "out_tok": 11930, "latency": 30.25, "id": "gen-1791297985-4fn0SdEQbD7hd9301l5v", "text": "<answer>3,87846,8,4,-586,-571</answer>"}
{"model": "google/gemini-3.8-flash", "task": "M1", "tier": "medium", "rep": 0, "pass": true, "detail": "3, 32, 10, 8, 2, 9", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.00460575, "in_tok": 271, "out_tok": 1174, "latency": 8.48, "id": "gen-1791298010-VDgKBH95aEQ3wsnX73yb", "text": "<answer>3, 32, 10, 8, 2, 9</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "M3", "tier": "medium", "rep": 0, "pass": true, "detail": "27 Dross 1130", "routed": "deepseek/deepseek-v4.1-flash", "provider": "CoreWeave", "cost": 0.00231325, "in_tok": 175, "out_tok": 3505, "latency": 22.17, "id": "gen-1791297999-yolfIOsF6H5xO0oApt89", "text": "<answer>27 Dross 1130</answer>"}
{"model": "google/gemini-3.8-flash", "task": "M2", "tier": "medium", "rep": 0, "pass": true, "detail": "37/37 hidden tests", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.00556125, "in_tok": 130, "out_tok": 1457, "latency": 10.68, "id": "gen-1791298011-afVGAnjvdU87EO9Azw2x", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    if not xs:\n        return []\n\n    n = len(xs)\n    out = [0] * n\n    stack = []\n\n    for i in range(n - 1, -1, -1):\n        while stack and xs[stack[-1]] < xs[i]:\n            stack.pop()\n        d = stack[-1] - i if stack else 0\n        out[i] = xs[i] + d\n        stack.append(i)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "google/gemini-3.8-flash", "task": "M3", "tier": "medium", "rep": 0, "pass": true, "detail": "27 Dross 1130", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.0061635, "in_tok": 163, "out_tok": 1611, "latency": 10.13, "id": "gen-1791298013-o8Hq2hqbIhIsrdV7S6Hn", "text": "<answer>27 Dross 1130</answer>"}
{"model": "google/gemini-3.8-flash", "task": "M4", "tier": "medium", "rep": 0, "pass": true, "detail": "grid ok", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.004284, "in_tok": 217, "out_tok": 1099, "latency": 8.67, "id": "gen-1791298015-cZYcuvH9jCKFhJICT3oP", "text": "<answer>{\"1\": {\"name\": \"Ilsa\", \"color\": \"slate\", \"drink\": \"mate\"}, \"2\": {\"name\": \"Bram\", \"color\": \"plum\", \"drink\": \"horchata\"}, \"3\": {\"name\": \"Dagny\", \"color\": \"olive\", \"drink\": \"chicha\"}, \"4\": {\"name\": \"Evander\", \"color\": \"teal\", \"drink\": \"kvass\"}}</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "E1", "tier": "easy", "rep": 0, "pass": true, "detail": "{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.000812, "in_tok": 136, "out_tok": 54, "latency": 2.64, "id": "gen-1791298023-6qE2sUuoPYONmOuinVum", "text": "<answer>{\"vendor\":\"Pellucid Ferry Co.\",\"total\":334,\"due\":\"2026-10-17\"}</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "E3", "tier": "easy", "rep": 0, "pass": true, "detail": "520", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.00054, "in_tok": 80, "out_tok": 38, "latency": 2.76, "id": "gen-1791298025-R9cfROHl46IXbkuiNs27", "text": "<answer>520</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "E4", "tier": "easy", "rep": 0, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.00069, "in_tok": 75, "out_tok": 54, "latency": 2.25, "id": "gen-1791298028-JzSlFHRUpyM0nOzwz0DZ", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "H2", "tier": "hard", "rep": 0, "pass": true, "detail": "46/46 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "InferenceNet", "cost": 0.0013696, "in_tok": 250, "out_tok": 4507, "latency": 28.51, "id": "gen-1791298002-HGWGkmCnReCWc2QfKxxT", "text": "```python\nimport heapq\n\ndef vault(grid: list[str]) -> int:\n    if not grid:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    start = end = None\n    portals = []\n\n    for r in range(rows):\n        for c, ch in enumerate(grid[r]):\n            if ch == 'S':\n                start = (r, c)\n            elif ch == 'E':\n                end = (r, c)\n            elif ch == '*':\n                portals.append((r, c))\n\n    if start is None or end is None:\n        return -1\n    if start == end:\n        return 0\n\n    portal_link = {}\n    if len(portals) == 2:\n        portal_link[portals[0]] = portals[1]\n        portal_link[portals[1]] = portals[0]\n\n    INF = 10**9\n    dist = [[[INF] * 8 for _ in range(cols)] for _ in range(rows)]\n    sr, sc = start\n    dist[sr][sc][0] = 0\n    heap = [(0, sr, sc, 0)]\n\n    directions = [(1, 0), (-1, 0), (0, 1), (0, -1)]\n\n    while heap:\n        cost, r, c, mask = heapq.heappop(heap)\n\n        if cost != dist[r][c][mask]:\n            continue\n        if (r, c) == end:\n            return cost\n\n        if (r, c) in portal_link:\n            nr, nc = portal_link[(r, c)]\n            if cost < dist[nr][nc][mask]:\n                dist[nr][nc][mask] = cost\n                heapq.heappush(heap, (cost, nr, nc, mask))\n\n        for dr, dc in directions:\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < rows and 0 <= nc < cols):\n                continue\n\n            ch = grid[nr][nc]\n            if ch == '#':\n                continue\n\n            if 'A' <= ch <= 'C':\n                needed = 1 << (ord(ch) - ord('A'))\n                if not (mask & needed):\n                    continue\n\n            next_mask = mask\n            if 'a' <= ch <= 'c':\n                next_mask |= 1 << (ord(ch) - ord('a'))\n\n            next_cost = cost + 1\n            if '1' <= ch <= '9':\n                next_cost += ord(ch) - ord('0')\n\n            if next_cost < dist[nr][nc][next_mask]:\n                dist[nr][nc][next_mask] = next_cost\n                heapq.heappush(heap, (next_cost, nr, nc, next_mask))\n\n    return -1\n```"}
{"model": "openai/gpt-6.1-sol", "task": "E2", "tier": "easy", "rep": 0, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.000524, "in_tok": 152, "out_tok": 22, "latency": 6.76, "id": "gen-1791298024-kRwOhDFah9Gg9vF3aIEl", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "H3", "tier": "hard", "rep": 0, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "deepseek/deepseek-v4.1-flash", "provider": "CoreWeave", "cost": 0.00478795, "in_tok": 244, "out_tok": 7291, "latency": 29.73, "id": "gen-1791298005-8LVOB58SyumVVGeXCLJ1", "text": "<answer>{\"J1\":0,\"J2\":8,\"J3\":0,\"J4\":5,\"J5\":6,\"J6\":1,\"J7\":7,\"J8\":11}</answer>"}
{"model": "google/gemini-3.8-flash", "task": "H1", "tier": "hard", "rep": 0, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.0181335, "in_tok": 328, "out_tok": 4770, "latency": 24.93, "id": "gen-1791298010-cKcKMik7nze40iGpZAqf", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "M2", "tier": "medium", "rep": 0, "pass": true, "detail": "37/37 hidden tests", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.001988, "in_tok": 124, "out_tok": 174, "latency": 4.95, "id": "gen-1791298031-uwtYTzCaVZRlwhHBs7G0", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n\n    out = xs.copy()\n    stack: list[int] = []\n\n    for j, value in enumerate(xs):\n        while stack and xs[stack[-1]] <= value:\n            i = stack.pop()\n            out[i] = xs[i] + (j - i)\n        stack.append(j)\n\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "openai/gpt-6.1-sol", "task": "M1", "tier": "medium", "rep": 0, "pass": true, "detail": "3, 32, 10, 8, 2, 9", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.002878, "in_tok": 264, "out_tok": 235, "latency": 7.2, "id": "gen-1791298030-vcUkRw7u3iWJRO5Lt0dy", "text": "<answer>3, 32, 10, 8, 2, 9</answer>"}
{"model": "google/gemini-3.8-flash", "task": "H2", "tier": "hard", "rep": 0, "pass": false, "detail": "4/46 hidden tests", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.00948, "in_tok": 225, "out_tok": 2483, "latency": 16.97, "id": "gen-1791298021-mbQhnIcTFP4kb8btG1JD", "text": "```python\nimport heapq\n\n\ndef vault(grid: list[str]) -> int:\n    R, C = len(grid), len(grid[0])\n    start, end = None, None\n    portals = []\n\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            if ch == \"S\":\n                start = (r, c)\n            elif ch == \"E\":\n                end = (r, c)\n            elif ch == \"*\":\n                portals.append((r, c))\n\n    if not start or not end:\n        return -1\n    if start == end:\n        return 0\n\n    has_portals = len(portals) == 2\n    p1, p2 = portals if has_portals else (None, None)\n\n    dist = {(start[0], start[1], 0): 0}\n    pq = [(0, start[0], start[1], 0)]\n\n    while pq:\n        d, r, c, keys = heapq.heappop(pq)\n\n        if (r, c) == end:\n            return d\n\n        if d > dist.get((r, c, keys), float(\"inf\")):\n            continue\n\n        # Portal jump\n        if has_portals and (r, c) in (p1, p2):\n            pr, pc = p2 if (r, c) == p1 else p1\n            if d < dist.get((pr, pc, keys), float(\"inf\")):\n                dist[(pr, pc, keys)] = d\n                heapq.heappush(pq, (d, pr, pc, keys))\n\n        # Normal moves\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n            ch = grid[nr][nc]\n            if ch == \"#\":\n                continue\n            if ch.isupper():\n                key_bit = 1 << (ord(ch.lower()) - ord(\"a\"))\n                if not (keys & key_bit):\n                    continue\n\n            cost = 1 + int(ch) if ch.isdigit() else 1\n            nkeys = keys | (1 << (ord(ch) - ord(\"a\"))) if ch.islower() else keys\n\n            if d + cost < dist.get((nr, nc, nkeys), float(\"inf\")):\n                dist[(nr, nc, nkeys)] = d + cost\n                heapq.heappush(pq, (d + cost, nr, nc, nkeys))\n\n    return -1\n```"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "H4", "tier": "hard", "rep": 0, "pass": true, "detail": "grid ok", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 0.004263417, "in_tok": 313, "out_tok": 7481, "latency": 41.69, "id": "gen-1791298001-Co78WtaE7R641qpicU79", "text": "<answer>{\"1\":{\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\"2\":{\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\"3\":{\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\"4\":{\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\"5\":{\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "M3", "tier": "medium", "rep": 0, "pass": true, "detail": "27 Dross 1130", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.00228, "in_tok": 150, "out_tok": 198, "latency": 7.92, "id": "gen-1791298035-vtlxYIA5caZ17iNA0DIE", "text": "<answer>27 Dross 1130</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "E2", "tier": "easy", "rep": 0, "pass": true, "detail": "TARN,WREN,GLINT,WREN,TARN,MOSS", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.001628, "in_tok": 232, "out_tok": 35, "latency": 3.01, "id": "gen-1791298043-B5PypjXkPjLxO6hdXL20", "text": "<answer>TARN,WREN,GLINT,WREN,TARN,MOSS</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "E1", "tier": "easy", "rep": 0, "pass": true, "detail": "{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.002276, "in_tok": 194, "out_tok": 75, "latency": 3.58, "id": "gen-1791298042-LZhL2a1rUccEhI7XpBjb", "text": "<answer>{\"vendor\": \"Pellucid Ferry Co.\", \"total\": 334, \"due\": \"2026-10-17\"}</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "M4", "tier": "medium", "rep": 0, "pass": true, "detail": "grid ok", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.002896, "in_tok": 218, "out_tok": 246, "latency": 11.44, "id": "gen-1791298039-jHFC8rNFzk5Qp9cqQnJx", "text": "<answer>{\"1\":{\"name\":\"Ilsa\",\"color\":\"slate\",\"drink\":\"mate\"},\"2\":{\"name\":\"Bram\",\"color\":\"plum\",\"drink\":\"horchata\"},\"3\":{\"name\":\"Dagny\",\"color\":\"olive\",\"drink\":\"chicha\"},\"4\":{\"name\":\"Evander\",\"color\":\"teal\",\"drink\":\"kvass\"}}</answer>"}
{"model": "google/gemini-3.8-flash", "task": "H3", "tier": "hard", "rep": 0, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.01344975, "in_tok": 218, "out_tok": 3543, "latency": 26.02, "id": "gen-1791298022-U9lmKa2POuM0b9N3UyLj", "text": "<answer>{\"J1\": 0, \"J2\": 6, \"J3\": 0, \"J4\": 5, \"J5\": 8, \"J6\": 1, \"J7\": 7, \"J8\": 11}</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "E3", "tier": "easy", "rep": 0, "pass": true, "detail": "520", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.00168, "in_tok": 115, "out_tok": 61, "latency": 3.74, "id": "gen-1791298046-ffn7IbAmYvP500w63Tr3", "text": "<answer>520</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "H3", "tier": "hard", "rep": 0, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.00436, "in_tok": 220, "out_tok": 392, "latency": 11.71, "id": "gen-1791298038-79BIupEkf2dTyQtFbZvg", "text": "<answer>{\"J1\":0,\"J2\":7,\"J3\":0,\"J4\":5,\"J5\":9,\"J6\":1,\"J7\":6,\"J8\":11}</answer>"}
{"model": "google/gemini-3.8-flash", "task": "H4", "tier": "hard", "rep": 0, "pass": true, "detail": "grid ok", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.0174045, "in_tok": 286, "out_tok": 4584, "latency": 31.86, "id": "gen-1791298018-HgSqib7WlUWve1KO1iKI", "text": "<answer>\n{\n  \"1\": {\n    \"name\": \"Fenna\",\n    \"color\": \"coral\",\n    \"drink\": \"chicha\"\n  },\n  \"2\": {\n    \"name\": \"Dagny\",\n    \"color\": \"teal\",\n    \"drink\": \"kvass\"\n  },\n  \"3\": {\n    \"name\": \"Bram\",\n    \"color\": \"slate\",\n    \"drink\": \"horchata\"\n  },\n  \"4\": {\n    \"name\": \"Evander\",\n    \"color\": \"ochre\",\n    \"drink\": \"sbiten\"\n  },\n  \"5\": {\n    \"name\": \"Gideon\",\n    \"color\": \"olive\",\n    \"drink\": \"mate\"\n  }\n}\n</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "H1", "tier": "hard", "rep": 0, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "deepseek/deepseek-v4.1-flash", "provider": "InferenceNet", "cost": 0.00383595, "in_tok": 345, "out_tok": 12706, "latency": 54.42, "id": "gen-1791297996-UL8XLVRq71u9XEgHVtNX", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "E4", "tier": "easy", "rep": 0, "pass": true, "detail": "tundra, lintel, saffron, ember, quiver, velvet", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.002616, "in_tok": 109, "out_tok": 109, "latency": 5.97, "id": "gen-1791298048-f7plOjWOCcJseuwe95pk", "text": "<answer>tundra, lintel, saffron, ember, quiver, velvet</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "M1", "tier": "medium", "rep": 0, "pass": true, "detail": "3,32,10,8,2,9", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.013048, "in_tok": 357, "out_tok": 581, "latency": 6.38, "id": "gen-1791298047-wOVeAmsb5LIeIFVBBQLw", "text": "Trace of the stack after each step (bottom to top):\n\n| Step | Stack |\n|---|---|\n| `3 4 8` | [3, 4, 8] |\n| `swap` | [3, 8, 4] |\n| `over` | [3, 8, 4, 8] |\n| `mul` (4\u00b78 = 32) | [3, 8, 32] |\n| `10 4 11` | [3, 8, 32, 10, 4, 11] |\n| `swap` | [3, 8, 32, 10, 11, 4] |\n| `sub` (11 \u2212 4 = 7) | [3, 8, 32, 10, 7] |\n| `4 mul` (7\u00b74 = 28) | [3, 8, 32, 10, 28] |\n| `mod` (10 mod 28 = 10) | [3, 8, 32, 10] |\n| `rot` (8 32 10 \u2192 32 10 8) | [3, 32, 10, 8] |\n| `2 9` | [3, 32, 10, 8, 2, 9] |\n\n<answer>3,32,10,8,2,9</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "H4", "tier": "hard", "rep": 0, "pass": true, "detail": "grid ok", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.008346, "in_tok": 283, "out_tok": 778, "latency": 17.83, "id": "gen-1791298036-mkUw3Vgqkb0Cd3qmvErQ", "text": "<answer>{\"1\":{\"name\":\"Fenna\",\"color\":\"coral\",\"drink\":\"chicha\"},\"2\":{\"name\":\"Dagny\",\"color\":\"teal\",\"drink\":\"kvass\"},\"3\":{\"name\":\"Bram\",\"color\":\"slate\",\"drink\":\"horchata\"},\"4\":{\"name\":\"Evander\",\"color\":\"ochre\",\"drink\":\"sbiten\"},\"5\":{\"name\":\"Gideon\",\"color\":\"olive\",\"drink\":\"mate\"}}</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "M2", "tier": "medium", "rep": 0, "pass": true, "detail": "37/37 hidden tests", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.006336, "in_tok": 179, "out_tok": 281, "latency": 5.05, "id": "gen-1791298049-F9BZij2zMPAWnGie5j1t", "text": "```python\ndef ripple(xs: list[int]) -> list[int]:\n    n = len(xs)\n    if n == 0:\n        return []\n    out = [0] * n\n    stack = []  # indices whose values decrease strictly from bottom to top\n    for i in range(n - 1, -1, -1):\n        while stack and xs[stack[-1]] < xs[i]:\n            stack.pop()\n        d = (stack[-1] - i) if stack else 0\n        out[i] = xs[i] + d\n        stack.append(i)\n    k = sum(xs) % n\n    return out[k:] + out[:k]\n```"}
{"model": "openai/gpt-6.1-sol", "task": "H1", "tier": "hard", "rep": 0, "pass": true, "detail": "3, 87846, 8, 4, -586, -571", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.012006, "in_tok": 318, "out_tok": 1137, "latency": 24.62, "id": "gen-1791298031-0uTcgQo0EnRMrFmsSu7i", "text": "<answer>3, 87846, 8, 4, -586, -571</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "M3", "tier": "medium", "rep": 0, "pass": true, "detail": "27 Dross 1130", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.018088, "in_tok": 207, "out_tok": 863, "latency": 9.18, "id": "gen-1791298050-CI3uN2L5xUdOgytotXvy", "text": "I counted the days left in 1122, then subtracted whole years until fewer than one year's days remained.\n\n**Year lengths**\n- A normal year has 31 + 28 + 30 + 33 + 29 + 31 + 30 = 212 days.\n- A leap year has 214 days, because Brine has 30 days.\n\n**Days left in 1122**\n- 1122 is not divisible by 5, so it is a normal year.\n- 16 Brine is day 31 + 16 = 47 of the year.\n- 212 \u2212 47 = 165 days take us to 30 Gale 1122.\n- That leaves 1769 \u2212 165 = 1604 days to add.\n\n**Whole years**\n\n| Year | Type | Days | Days left after the year |\n|------|------|------|--------------------------|\n| 1123 | normal | 212 | 1392 |\n| 1124 | normal | 212 | 1180 |\n| 1125 | leap (divisible by 5, not by 40) | 214 | 966 |\n| 1126 | normal | 212 | 754 |\n| 1127 | normal | 212 | 542 |\n| 1128 | normal | 212 | 330 |\n| 1129 | normal | 212 | 118 |\n\n**Day 118 of 1130**\n- 1130 is a leap year, since it is divisible by 5 but not by 40, so Brine has 30 days.\n- After Ashen (31 days), 118 \u2212 31 = 87 days remain.\n- After Brine (30 days), 87 \u2212 30 = 57 days remain.\n- After Cinder (30 days), 57 \u2212 30 = 27 days remain.\n- The result is day 27 of Dross.\n\n<answer>27 Dross 1130</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "M4", "tier": "medium", "rep": 0, "pass": true, "detail": "grid ok", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.01884, "in_tok": 315, "out_tok": 879, "latency": 9.25, "id": "gen-1791298050-lBgezOD6tvgkMvEtjrwn", "text": "Working through the clues in order of how much they fix:\n\n1. **Kvass is in house 4.** Clue 5 says plum is left of kvass, so kvass is in house 2, 3 or 4. Clue 1 says Ilsa is left of plum, so plum is in house 2 or later, which pushes kvass to house 3 or 4. House 3 is chicha (clue 7), so kvass is in house 4.\n2. **Olive is in house 3.** Clue 4 puts olive next to the kvass drinker in house 4.\n3. **Plum is in house 2, and Ilsa is in house 1.** Plum must be left of house 4 and is not house 3, so it is house 2. Ilsa must be left of plum, so she is in house 1.\n4. **Slate is house 1, teal is house 4.** Slate and teal fill houses 1 and 4. Dagny must be left of teal (clue 6), so teal cannot be house 1.\n5. **Bram is in house 2.** Clue 3 puts Bram next to slate (house 1), so he is in house 2.\n6. **Mate is house 1, horchata is house 2.** Clue 2 puts mate next to Bram, so in house 1 or 3. House 3 is chicha, so mate is house 1, leaving horchata for house 2.\n7. **Dagny is house 3, Evander is house 4.** Dagny must be left of teal (house 4), and houses 1 and 2 are taken, so she is in house 3.\n\n<answer>{\"1\": {\"name\": \"Ilsa\", \"color\": \"slate\", \"drink\": \"mate\"}, \"2\": {\"name\": \"Bram\", \"color\": \"plum\", \"drink\": \"horchata\"}, \"3\": {\"name\": \"Dagny\", \"color\": \"olive\", \"drink\": \"chicha\"}, \"4\": {\"name\": \"Evander\", \"color\": \"teal\", \"drink\": \"kvass\"}}</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "H2", "tier": "hard", "rep": 0, "pass": true, "detail": "46/46 hidden tests", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.006982, "in_tok": 221, "out_tok": 654, "latency": 21.65, "id": "gen-1791298038-m7pX2zMEA2m6RaeifX4g", "text": "```python\ndef vault(grid: list[str]) -> int:\n    from heapq import heappop, heappush\n\n    if not grid or not grid[0]:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    start = None\n    portals = []\n\n    for r in range(rows):\n        for c in range(cols):\n            if grid[r][c] == \"S\":\n                start = (r, c)\n            elif grid[r][c] == \"*\":\n                portals.append((r, c))\n\n    if start is None:\n        return -1\n\n    links = {}\n    if len(portals) == 2:\n        a, b = portals\n        links[a] = b\n        links[b] = a\n\n    distances = [[[float(\"inf\")] * 8 for _ in range(cols)]\n                 for _ in range(rows)]\n    sr, sc = start\n    distances[sr][sc][0] = 0\n    heap = [(0, sr, sc, 0)]\n\n    while heap:\n        cost, r, c, keys = heappop(heap)\n        if cost != distances[r][c][keys]:\n            continue\n        if grid[r][c] == \"E\":\n            return cost\n\n        if (r, c) in links:\n            nr, nc = links[(r, c)]\n            if cost < distances[nr][nc][keys]:\n                distances[nr][nc][keys] = cost\n                heappush(heap, (cost, nr, nc, keys))\n\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < rows and 0 <= nc < cols):\n                continue\n\n            cell = grid[nr][nc]\n            if cell == \"#\":\n                continue\n            if cell in \"ABC\" and not (keys & (1 << (ord(cell) - ord(\"A\")))):\n                continue\n\n            next_keys = keys\n            if cell in \"abc\":\n                next_keys |= 1 << (ord(cell) - ord(\"a\"))\n\n            next_cost = cost + 1\n            if cell in \"123456789\":\n                next_cost += int(cell)\n\n            if next_cost < distances[nr][nc][next_keys]:\n                distances[nr][nc][next_keys] = next_cost\n                heappush(heap, (next_cost, nr, nc, next_keys))\n\n    return -1\n```"}
{"model": "anthropic/claude-opus-5.5", "task": "H2", "tier": "hard", "rep": 0, "pass": true, "detail": "46/46 hidden tests", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.020024, "in_tok": 321, "out_tok": 937, "latency": 9.05, "id": "gen-1791298052-rNUHCPhOpjqoYkWUmNHv", "text": "```python\nimport heapq\n\ndef vault(grid: list[str]) -> int:\n    if not grid:\n        return -1\n    R = len(grid)\n    C = len(grid[0]) if R else 0\n    start = None\n    portals = []\n    for r in range(R):\n        for c in range(len(grid[r])):\n            ch = grid[r][c]\n            if ch == 'S':\n                start = (r, c)\n            elif ch == '*':\n                portals.append((r, c))\n    if start is None:\n        return -1\n    link = {}\n    if len(portals) == 2:\n        link[portals[0]] = portals[1]\n        link[portals[1]] = portals[0]\n\n    def cell(r, c):\n        if 0 <= r < R and 0 <= c < len(grid[r]):\n            return grid[r][c]\n        return '#'\n\n    INF = float('inf')\n    dist = {}\n    s = (start[0], start[1], 0)\n    dist[s] = 0\n    pq = [(0, start[0], start[1], 0)]\n    while pq:\n        d, r, c, m = heapq.heappop(pq)\n        if dist.get((r, c, m), INF) < d:\n            continue\n        if grid[r][c] == 'E':\n            return d\n        moves = []\n        for dr, dc in ((1, 0), (-1, 0), (0, 1), (0, -1)):\n            nr, nc = r + dr, c + dc\n            ch = cell(nr, nc)\n            if ch == '#':\n                continue\n            if ch in 'ABC':\n                if not (m >> (ord(ch) - ord('A'))) & 1:\n                    continue\n            cost = 1\n            if ch.isdigit():\n                cost += int(ch)\n            nm = m\n            if ch in 'abc':\n                nm = m | (1 << (ord(ch) - ord('a')))\n            moves.append((nr, nc, nm, cost))\n        if (r, c) in link:\n            pr, pc = link[(r, c)]\n            moves.append((pr, pc, m, 0))\n        for nr, nc, nm, cost in moves:\n            nd = d + cost\n            key = (nr, nc, nm)\n            if nd < dist.get(key, INF):\n                dist[key] = nd\n                heapq.heappush(pq, (nd, nr, nc, nm))\n    return -1\n```"}
{"model": "anthropic/claude-opus-5.5", "task": "H1", "tier": "hard", "rep": 0, "pass": true, "detail": "3,87846,8,4,-586,-571", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.040092, "in_tok": 453, "out_tok": 1914, "latency": 15.32, "id": "gen-1791298048-BbFqR7ikWMqBnmNAfxsx", "text": "Simulating the program token by token (stack shown bottom to top):\n\n1. `8 3 6` gives [8, 3, 6].\n2. `rep 4 [ 11 mul ]` gives 6\u00b711\u2074 = 87846, so [8, 3, 87846].\n3. `rot` gives [3, 87846, 8].\n4. `10` gives [3, 87846, 8, 10].\n5. `rep 4 [ 10 add 4 5 ]`: each pass adds 10 to the top, then pushes 4 and 5.\n   - Pass 1: 10 becomes 20, giving [3, 87846, 8, 20, 4, 5].\n   - Passes 2 to 4: the top 5 becomes 15 each time, then 4 and 5 are pushed.\n   - Result: [3, 87846, 8, 20, 4, 15, 4, 15, 4, 15, 4, 5].\n6. `over` gives [3, 87846, 8, 20, 4, 15, 4, 15, 4, 15, 4, 5, 4].\n7. `rep 3 [ drop zig ]`:\n   - Pass 1: drop 4, then zig 5 gives 16.\n   - Pass 2: drop 16, then zig 4 gives 2.\n   - Pass 3: drop 2, then zig 15 gives 46.\n   - Result: [3, 87846, 8, 20, 4, 15, 4, 15, 4, 46].\n8. `6 add` gives 52, so [3, 87846, 8, 20, 4, 15, 4, 15, 4, 52].\n9. `over 3` gives [3, 87846, 8, 20, 4, 15, 4, 15, 4, 52, 4, 3].\n10. `rep 2 [ mul 2 drop ]`:\n    - Pass 1: 4\u00b73 = 12.\n    - Pass 2: 52\u00b712 = 624.\n    - Result: [3, 87846, 8, 20, 4, 15, 4, 15, 4, 624].\n11. `sub`: 4 \u2212 624 = \u2212620, giving [3, 87846, 8, 20, 4, 15, 4, 15, \u2212620].\n12. `2 ifz [ ... ]`: pops 2, which is not zero, so the body is skipped.\n13. `add`: 15 + (\u2212620) = \u2212605, giving [3, 87846, 8, 20, 4, 15, 4, \u2212605].\n14. `mod`: 4 % \u2212605 = \u2212601, giving [3, 87846, 8, 20, 4, 15, \u2212601].\n15. `mod`: 15 % \u2212601 = \u2212586, giving [3, 87846, 8, 20, 4, \u2212586].\n16. `rot` gives [3, 87846, 8, 4, \u2212586, 20].\n17. `5 sub`: 20 \u2212 5 = 15, giving [3, 87846, 8, 4, \u2212586, 15].\n18. `over` gives [3, 87846, 8, 4, \u2212586, 15, \u2212586].\n19. `mod`: 15 % \u2212586 = \u2212571, giving [3, 87846, 8, 4, \u2212586, \u2212571].\n20. `1 ifz [ ... ]`: pops 1, which is not zero, so the body is skipped.\n\n<answer>3,87846,8,4,-586,-571</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "H3", "tier": "hard", "rep": 0, "pass": true, "detail": "makespan 16 (optimal 16)", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.023748, "in_tok": 292, "out_tok": 1129, "latency": 11.54, "id": "gen-1791298053-DjoNDn7YJ9Mdkw5U4JDw", "text": "This schedule reaches a makespan of 16, which is the minimum.\n\n**Lower bound**\n- J8 must follow J3, J4, J5, J6 and J7 directly. J1 must finish before J4, so J1 also precedes J8 indirectly.\n- J2 is hot, so it cannot run at the same time as J8. It must either finish before J8 starts or run after J8 finishes.\n- If J2 runs after J8, the makespan becomes J8's finish time plus 2.\n- If J2 runs before J8, then all seven other jobs must finish before J8 starts. Their total work is 5+2+1+1+2+6+4 = 21. Split over 2 machines, that needs at least 10.5 time units, so at least 11 with integer times.\n- J8 therefore starts at 11 at the earliest and finishes at 16 at the earliest. So the makespan is at least 16.\n\n**Schedule achieving 16**\n\n| Machine | Jobs (start\u2013finish) |\n|---|---|\n| A | J3 (0\u20131), J6 (1\u20137), J7 (7\u201311), J8 (11\u201316) |\n| B | J1 (0\u20135), J4 (5\u20136), J2 (6\u20138), J5 (8\u201310) |\n\n**Checks**\n- Precedence:\n  - J1 finishes at 5, before J4 starts at 5.\n  - J3 finishes at 1, before J6 starts at 1.\n  - J3, J4, J5, J6 and J7 all finish by 11, which is when J8 starts.\n- Heat rule:\n  - J2 runs 6\u20138 and J5 runs 8\u201310, so they do not overlap.\n  - J8 runs 11\u201316, after both of them have finished.\n\n<answer>{\"J1\": 0, \"J2\": 6, \"J3\": 0, \"J4\": 5, \"J5\": 8, \"J6\": 1, \"J7\": 7, \"J8\": 11}</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "H4", "tier": "hard", "rep": 0, "pass": true, "detail": "grid ok", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.044028, "in_tok": 402, "out_tok": 2121, "latency": 19.56, "id": "gen-1791298051-cWqCtvIjZHfBgtitVGig", "text": "Only one arrangement satisfies all 11 clues, found by placing Dagny first and testing each case.\n\n**Placing Dagny and Fenna (clues 4, 8, 9, 10, 5)**\n- Fenna drinks chicha and house 2 drinks kvass, so Fenna cannot be in house 2.\n- Dagny cannot be in house 5, because she must be left of Bram.\n- **Dagny in 1:** Fenna would have to be in 2. Impossible.\n- **Dagny in 3:** Fenna must be in 4 and Bram in 5. Gideon and Evander then take houses 1 and 2. Slate must be left of Gideon, so slate is 1, Gideon is 2 and ochre is 2. Evander is in 1, so mate must be in 2, but house 2 drinks kvass. Impossible.\n- **Dagny in 4:** Bram is in 5 and Fenna in 3, so Gideon is in 1 or 2. The slate\u2013ochre pair can sit at (1,2), (2,3) or (3,2).\n  - In every placement, either slate is not left of Gideon,\n  - or Evander (in house 1) cannot be next to the mate drinker. Impossible.\n- **Dagny in 2 (teal):** Fenna is in 1 or 3.\n\n**Case: Fenna in 1, drinking chicha**\n- Bram is in 3, 4 or 5, and the slate\u2013ochre pair lies within houses 3\u20135.\n- Olive cannot be in house 1, since something must be immediately left of it (clue 3). So house 1 is coral, and houses 3\u20135 are olive, slate and ochre.\n- **Olive in 3:** sbiten would be in 2, which is kvass. Impossible.\n- **Ochre 3, slate 4, olive 5:** Gideon must be in 5 and sbiten in 4, so mate is in 5. Evander is then in 4, leaving Bram in the ochre house 3. This breaks clue 6.\n- **Slate 3, ochre 4, olive 5:**\n  - Clue 3 puts sbiten in 4, and clue 7 then puts mate in 5. That leaves horchata for 3.\n  - Evander must be next to mate, so he is in 4. Gideon (right of slate) is in 5, and Bram is in 3.\n  - Every clue checks: Bram is not ochre, and Dagny (2) is left of Bram (3).\n\nThe Fenna-in-3 branch was not fully checked; the puzzle states the solution is unique, so this valid arrangement is taken as the answer.\n\n<answer>{\"1\": {\"name\": \"Fenna\", \"color\": \"coral\", \"drink\": \"chicha\"}, \"2\": {\"name\": \"Dagny\", \"color\": \"teal\", \"drink\": \"kvass\"}, \"3\": {\"name\": \"Bram\", \"color\": \"slate\", \"drink\": \"horchata\"}, \"4\": {\"name\": \"Evander\", \"color\": \"ochre\", \"drink\": \"sbiten\"}, \"5\": {\"name\": \"Gideon\", \"color\": \"olive\", \"drink\": \"mate\"}}</answer>"}
{"model": "typesafe/jev-router", "task": "X2", "tier": "extreme", "rep": 1, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.002922, "in_tok": 201, "out_tok": 252, "latency": 8.17, "id": "gen-1791298176-uK2kvVqxf9nRCKr3zobz", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "typesafe/jev-router", "task": "X2", "tier": "extreme", "rep": 2, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.013132, "in_tok": 291, "out_tok": 1255, "latency": 10.05, "id": "gen-1791298176-w9urjYmqvcwzxL6GFzNL", "text": "Adding 42453 days to 22 Gale 1228 gives Wick 5 Gale 1428.\n\n**Weekday.** The week has 9 days, and 42453 mod 9 = 0 (its digits sum to 18). The weekday is therefore unchanged: Wick.\n\n**Date.**\n- A normal year has 31+28+30+33+29+31+30 = 212 days, and a leap year has 214.\n- 1228 is not divisible by 5, so it is a normal year. From 22 Gale 1228, there are 8 days to 30 Gale and 1 more to 1 Ashen 1229, which is 9 days in total.\n- The remaining offset is 42453 \u2212 9 = 42444 days from 1 Ashen 1229.\n- In any 40-year span, 8 years are divisible by 5 and 1 of those is divisible by 40, so there are 7 leap years. That gives 40\u00b7212 + 7\u00b72 = 8494 days per 40 years.\n- The 200 years from 1229 to 1428 contain 35 leap years, so they total 200\u00b7212 + 70 = 42470 days. This puts 1 Ashen 1429 at an offset of 42470.\n- Our offset is 42444, which is 26 days before 1 Ashen 1429. That falls in 1428, a normal year of 212 days.\n- Counting back, the day is the 212 \u2212 26 + 1 = 187th day of 1428.\n- Cumulative days at the end of each month are Ashen 31, Brine 59, Cinder 89, Dross 122, Ember 151, Fallow 182. Day 187 is therefore day 5 of Gale.\n\n<answer>Wick 5 Gale 1428</answer>"}
{"model": "typesafe/jev-router", "task": "X4", "tier": "extreme", "rep": 0, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.014956, "in_tok": 343, "out_tok": 1427, "latency": 10.47, "id": "gen-1791298176-07IVdwikydNhCqqvaKS3", "text": "This schedule finishes at time 21. I found it by working out the critical chain, so I have not run an exhaustive search; the argument below shows why 21 is the best possible.\n\n**Why 21 is the minimum**\n\n- J1 cannot start before 4. J5 cannot start before 6, so J5 cannot finish before 10.\n- J1 and J5 are both hot, so they cannot overlap. Because J5 cannot finish before 10, J5 cannot run before J1 unless J1 starts at 10 or later.\n- J8 must follow both J1 and J5, and J9 must follow J8.\n- If J1 runs first (4\u201310) and J5 follows (10\u201314), then J8 runs 14\u201318 and J9 runs 18\u201321. The makespan is 21.\n- If J5 runs first (6\u201310), J1 runs 10\u201316, J8 runs 16\u201320 and J9 runs 20\u201323. The makespan is 23, which is worse.\n\n**Schedule**\n\n- Machine 1: J1 (4\u201310), J5 (10\u201314), J8 (14\u201318), J9 (18\u201321).\n- Machine 2: J2 (0\u20131), J4 (1\u20132), J6 (2\u20137), J3 (10\u201312), J7 (14\u201316).\n\n**Checks**\n\n- Release times hold: J1 starts at 4, J3 at 10, J5 at 10 and J8 at 14.\n- Precedences hold:\n  - J1 ends at 10, before J3 (10) and J8 (14).\n  - J2 ends at 1, before J3 (10).\n  - J4 ends at 2, J5 ends at 14 and J6 ends at 7, so all finish by the time J7 starts at 14.\n  - J5 ends at 14, before J8 (14).\n  - J8 ends at 18, before J9 (18).\n- The hot jobs J1 (4\u201310), J2 (0\u20131) and J5 (10\u201314) never overlap.\n\n<answer>{\"J1\": 4, \"J2\": 0, \"J3\": 10, \"J4\": 1, \"J5\": 10, \"J6\": 2, \"J7\": 14, \"J8\": 14, \"J9\": 18}</answer>"}
{"model": "typesafe/jev-router", "task": "X4", "tier": "extreme", "rep": 1, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.013136, "in_tok": 343, "out_tok": 1245, "latency": 11.65, "id": "gen-1791298176-IlCB8ENvVeTK0bjRrYrV", "text": "This schedule finishes at time 21. I did not run a formal optimality proof; the bound below is a hand argument.\n\n**Why 21 is the best I found**\n\n- J1 cannot start before 4, so it finishes at 10 at the earliest.\n- J5 cannot start before 6, so without a clash it would run 6\u201310. The heat rule forbids J5 overlapping J1, so one of them must run entirely before the other.\n- If J5 runs first (6\u201310), then J1 runs 10\u201316, J8 runs 16\u201320 and J9 runs 20\u201323, so the makespan is 23.\n- If J1 runs first (4\u201310), then J5 runs 10\u201314, J8 runs 14\u201318 and J9 runs 18\u201321, so the makespan is 21. This is the better option, and the chain J1 \u2192 J5 \u2192 J8 \u2192 J9 fixes the makespan at 21.\n\n**Placing the remaining jobs on the second machine**\n\n- J2 runs 0\u20131, which is before J1 starts, so the heat rule holds.\n- J6 runs 1\u20136 and J4 runs 6\u20137.\n- J3 needs both J1 and J2 finished, so it runs 10\u201312.\n- J7 needs J4, J5 and J6 finished, so it runs 14\u201316.\n\n**Machine assignment**\n\n- Machine 1: J1 (4\u201310), J5 (10\u201314), J8 (14\u201318), J9 (18\u201321).\n- Machine 2: J2 (0\u20131), J6 (1\u20136), J4 (6\u20137), J3 (10\u201312), J7 (14\u201316).\n\nAll release times, precedence constraints and the heat rule are satisfied.\n\n<answer>{\"J1\":4,\"J2\":0,\"J3\":10,\"J4\":6,\"J5\":10,\"J6\":1,\"J7\":14,\"J8\":14,\"J9\":18}</answer>"}
{"model": "typesafe/jev-router", "task": "X2", "tier": "extreme", "rep": 0, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.016632, "in_tok": 291, "out_tok": 1605, "latency": 12.35, "id": "gen-1791298176-2lYbDEVjziOGiOraAWrd", "text": "Starting from 22 Gale 1228, the date 42453 days later is Wick 5 Gale 1428. The working is below.\n\n**Year lengths.** A normal year has 31+28+30+33+29+31+30 = 212 days, and a leap year has 214 days. A leap year is a multiple of 5 that is not a multiple of 40. Any 40 consecutive years therefore contain 8 \u2212 1 = 7 leap years and total 40\u00b7212 + 7\u00b72 = 8494 days.\n\n**Reaching the start of 1229.** From 22 Gale 1228 there are 8 days left in 1228, so 1 Ashen 1229 is 9 days after the start date. The target is therefore 42453 \u2212 9 = 42444 days after 1 Ashen 1229.\n\n**Whole years.**\n- Four 40-year blocks cover 1229\u20131388 and use 4\u00b78494 = 33976 days. That leaves 42444 \u2212 33976 = 8468 days, counted from 1 Ashen 1389.\n- The 39 years 1389\u20131427 contain 7 leap years (1390, 1395, 1405, 1410, 1415, 1420, 1425; 1400 is a multiple of 40, so it is excluded). They total 39\u00b7212 + 14 = 8282 days.\n- That leaves 8468 \u2212 8282 = 186 days into 1428. Year 1428 is not a leap year.\n\n**Month within 1428.** Counting from 0 on 1 Ashen, the cumulative days at the end of each month are Ashen 31, Brine 59, Cinder 89, Dross 122, Ember 151, Fallow 182. Offset 186 is 4 days past 182, so the date is 5 Gale 1428.\n\n**Weekday.** 42453 = 9 \u00b7 4717, so the shift is 0 mod 9 and the weekday is the same as the start, Wick.\n\n<answer>Wick 5 Gale 1428</answer>"}
{"model": "typesafe/jev-router", "task": "X3", "tier": "extreme", "rep": 0, "pass": true, "detail": "66/66 hidden tests", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.010622, "in_tok": 276, "out_tok": 1007, "latency": 20.23, "id": "gen-1791298176-bNYEJ7TeTq83LWkWSKm1", "text": "```python\ndef vault2(grid: list[str]) -> int:\n    from heapq import heappop, heappush\n\n    if not grid or not grid[0]:\n        return -1\n\n    height, width = len(grid), len(grid[0])\n    cells = \"\".join(grid)\n    size = len(cells)\n    start = cells.find(\"S\")\n    end = cells.find(\"E\")\n    if start == -1 or end == -1:\n        return -1\n\n    costs = [1] * size\n    key_bits = [0] * size\n    door_bits = [0] * size\n    door_letters = [-1] * size\n    keys_by_letter = [0] * 3\n    doors_by_letter = [0] * 3\n    key_count = door_count = 0\n\n    for pos, cell in enumerate(cells):\n        if \"1\" <= cell <= \"9\":\n            costs[pos] += int(cell)\n        elif cell in \"abc\":\n            letter = ord(cell) - ord(\"a\")\n            bit = 1 << key_count\n            key_count += 1\n            key_bits[pos] = bit\n            keys_by_letter[letter] |= bit\n        elif cell in \"ABC\":\n            letter = ord(cell) - ord(\"A\")\n            bit = 1 << door_count\n            door_count += 1\n            door_bits[pos] = bit\n            door_letters[pos] = letter\n            doors_by_letter[letter] |= bit\n\n    neighbors = [[] for _ in range(size)]\n    for pos, cell in enumerate(cells):\n        if cell == \"#\":\n            continue\n        row, col = divmod(pos, width)\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = row + dr, col + dc\n            if 0 <= nr < height and 0 <= nc < width:\n                nxt = nr * width + nc\n                if cells[nxt] != \"#\":\n                    neighbors[pos].append(nxt)\n\n    # Carried keys equal collected key cells minus opened door cells,\n    # counted separately for each letter.\n    distances = {(start, 0, 0): 0}\n    heap = [(0, start, 0, 0)]\n\n    while heap:\n        cost, pos, picked, opened = heappop(heap)\n        if distances.get((pos, picked, opened)) != cost:\n            continue\n        if pos == end:\n            return cost\n\n        for nxt in neighbors[pos]:\n            next_picked = picked | key_bits[nxt]\n            next_opened = opened\n            door_bit = door_bits[nxt]\n\n            if door_bit and not (opened & door_bit):\n                letter = door_letters[nxt]\n                collected = (picked & keys_by_letter[letter]).bit_count()\n                consumed = (opened & doors_by_letter[letter]).bit_count()\n                if collected <= consumed:\n                    continue\n                next_opened |= door_bit\n\n            next_cost = cost + costs[nxt]\n            state = (nxt, next_picked, next_opened)\n            previous = distances.get(state)\n            if previous is None or next_cost < previous:\n                distances[state] = next_cost\n                heappush(heap, (next_cost, nxt, next_picked, next_opened))\n\n    return -1\n```"}
{"model": "typesafe/jev-router", "task": "X4", "tier": "extreme", "rep": 2, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.014036, "in_tok": 343, "out_tok": 1335, "latency": 11.53, "id": "gen-1791298186-BQdcKi5pTIt2kkJdlB88", "text": "Makespan 21 is the best I could find: the argument below shows no schedule can start J8 before time 14, and the schedule given achieves 21, but I did not prove 21 is optimal.\n\n**Why the heat rule forces a long critical path**\n- J1 cannot start before time 4, so it finishes at 10 at the earliest.\n- J5 cannot start before time 6, so it cannot run before J1 starts. If J5 ran in [6,10] it would overlap J1 when J1 starts at 4, which the heat rule forbids.\n- Delaying J1 until after J5 pushes J1 to finish at 16 or later, which is worse.\n- So the best order is J1 in [4,10], then J5 in [10,14].\n- J8 must follow both J1 and J5, so it runs [14,18]. J9 follows J8 and runs [18,21].\n\n**Fitting the other jobs**\n- Machine 1 runs the critical chain: J1 [4,10], J5 [10,14], J8 [14,18], J9 [18,21].\n- Machine 2 runs the rest:\n  - J2 runs [0,1], before J1 starts, so there is no hot overlap.\n  - J4 runs [1,2] and J6 runs [2,7].\n  - J3 runs [10,12], after both J1 and J2 have finished and after its release time of 4.\n  - J7 runs [14,16], after J4, J5 and J6 have finished.\n\nAll release times, precedence constraints and the heat rule are satisfied.\n\n<answer>{\"J1\": 4, \"J2\": 0, \"J3\": 10, \"J4\": 1, \"J5\": 10, \"J6\": 2, \"J7\": 14, \"J8\": 14, \"J9\": 18}</answer>"}
{"model": "typesafe/jev-router", "task": "X1", "tier": "extreme", "rep": 2, "pass": true, "detail": "1, 4, 4, -132, 11, 7", "routed": "anthropic/claude-sonnet-5.5", "provider": "Claude Platform on AWS", "cost": 0.036696, "in_tok": 518, "out_tok": 3566, "latency": 23.04, "id": "gen-1791298176-DAzHp4D1TrTtgQ53Tqd5", "text": "Final stack, bottom to top: 1, 4, 4, \u2212132, 11, 7.\n\nI traced the program step by step, tracking the full stack after each token group.\n\n1. `4 8 3 mod` gives [4, 2]. `zig` turns 2 into 1, so [4, 1].\n2. `0 ifz [ 1 ]` pops 0, runs the body and pushes 1, so [4, 1, 1]. `rot` gives [1, 1, 4].\n3. `rep 2 [ swap 4 ]` gives [1, 4, 4, 1, 4]. `over` and `7` give [1, 4, 4, 1, 4, 1, 7].\n4. `rep 3 [ over 2 sub swap ]` gives [1, 4, 4, 1, 4, 1, \u22121, \u22123, \u22125, 7].\n5. `dup 7 9 1` gives [..., 7, 7, 7, 9, 1]. `ifz` pops 1, which is nonzero, so the body is skipped.\n6. `mod` computes 7 % 9 = 7, then `sub` computes 7 \u2212 7 = 0. `over mod` computes 0 % 7 = 0. The stack is [1, 4, 4, 1, 4, 1, \u22121, \u22123, \u22125, 7, 0].\n7. `rep 3 [ 9 7 1 mul ]` pushes 9, 7 three times, giving [..., 0, 9, 7, 9, 7, 9, 7]. `rot` makes the tail 0, 9, 7, 9, 9, 7, 7.\n8. `rep 2 [ 3 1 add ]` pushes 4, 4. `add` makes 8, and `swap` gives [..., 9, 9, 7, 8, 7].\n9. `zig` turns 7 into 22, and `drop` removes it, leaving the tail 0, 9, 7, 9, 9, 7, 8.\n10. `8 3 3 mul` pushes 8 and 9. `zig` turns 9 into 28. `sub` gives 8 \u2212 28 = \u221220, and the next `sub` gives 8 + 20 = 28.\n11. `mod` computes 7 % 28 = 7. `9 sub` gives \u22122, and `add` gives 9 + (\u22122) = 7. The tail is now 0, 9, 7, 9, 7.\n12. `4 drop drop` pushes 4, then removes 4 and 7, leaving 0, 9, 7, 9. `add` gives 0, 9, 16, and `sub` gives 9 \u2212 16 = \u22127. The stack is [1, 4, 4, 1, 4, 1, \u22121, \u22123, \u22125, 7, 0, \u22127].\n13. `swap over` gives [..., 7, \u22127, 0, \u22127].\n14. `rep 4 [ add add dup drop ]` (the `dup drop` does nothing) has four iterations, each of which adds twice:\n    - Iteration 1 ends at [..., \u22125, 7, \u221214].\n    - Iteration 2 ends at [..., \u22121, \u22123, \u221212].\n    - Iteration 3 ends at [..., 1, \u221216].\n    - Iteration 4 ends at [1, 4, 4, 1, \u221211].\n15. `12 rot` gives [1, 4, 4, \u221211, 12, 1]. `mul` gives 12, then `mul` gives \u2212132, so [1, 4, 4, \u2212132]. `11 7` pushes the last two values.\n\n<answer>1, 4, 4, -132, 11, 7</answer>"}
{"model": "typesafe/jev-router", "task": "X1", "tier": "extreme", "rep": 0, "pass": true, "detail": "1, 4, 4, -132, 11, 7", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.013368, "in_tok": 364, "out_tok": 1264, "latency": 23.29, "id": "gen-1791298176-BHLFy1uDkog16qEnBvZt", "text": "<answer>1, 4, 4, -132, 11, 7</answer>"}
{"model": "typesafe/jev-router", "task": "X3", "tier": "extreme", "rep": 1, "pass": true, "detail": "66/66 hidden tests", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.012272, "in_tok": 276, "out_tok": 1172, "latency": 24.49, "id": "gen-1791298176-LHxdzhawRea92Pccs7Un", "text": "```python\ndef vault2(grid: list[str]) -> int:\n    from heapq import heappop, heappush\n\n    if not grid or not grid[0]:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    cells = \"\".join(grid)\n    n = rows * cols\n\n    start = cells.find(\"S\")\n    end = cells.find(\"E\")\n    if start == -1 or end == -1:\n        return -1\n\n    key_bits = [0] * n\n    door_bits = [0] * n\n    door_letters = [-1] * n\n    effects = []\n    costs = [1] * n\n\n    for pos, cell in enumerate(cells):\n        if \"1\" <= cell <= \"9\":\n            costs[pos] += int(cell)\n        elif cell in \"abc\":\n            key_bits[pos] = 1 << len(effects)\n            effects.append((ord(cell) - ord(\"a\"), 1))\n        elif cell in \"ABC\":\n            door_bits[pos] = 1 << len(effects)\n            letter = ord(cell) - ord(\"A\")\n            door_letters[pos] = letter\n            effects.append((letter, -1))\n\n    # Each bit records a collected key cell or an opened door cell.\n    # Thus the carried inventory is determined entirely by this mask.\n    mask_count = 1 << len(effects)\n    inventory = [(0, 0, 0)] * mask_count\n    for mask in range(1, mask_count):\n        bit = mask & -mask\n        previous = inventory[mask ^ bit]\n        letter, change = effects[bit.bit_length() - 1]\n        counts = list(previous)\n        counts[letter] += change\n        inventory[mask] = tuple(counts)\n\n    neighbors = [[] for _ in range(n)]\n    for pos, cell in enumerate(cells):\n        if cell == \"#\":\n            continue\n        row, col = divmod(pos, cols)\n        if row > 0 and cells[pos - cols] != \"#\":\n            neighbors[pos].append(pos - cols)\n        if row + 1 < rows and cells[pos + cols] != \"#\":\n            neighbors[pos].append(pos + cols)\n        if col > 0 and cells[pos - 1] != \"#\":\n            neighbors[pos].append(pos - 1)\n        if col + 1 < cols and cells[pos + 1] != \"#\":\n            neighbors[pos].append(pos + 1)\n\n    # State encoding: mask * n + position.\n    distances = [float(\"inf\")] * (mask_count * n)\n    distances[start] = 0\n    heap = [(0, start)]\n\n    while heap:\n        distance, state = heappop(heap)\n        if distance != distances[state]:\n            continue\n\n        mask, pos = divmod(state, n)\n        if pos == end:\n            return distance\n\n        for nxt in neighbors[pos]:\n            next_mask = mask\n            door_bit = door_bits[nxt]\n            if door_bit and not (mask & door_bit):\n                if inventory[mask][door_letters[nxt]] <= 0:\n                    continue\n                next_mask |= door_bit\n\n            next_mask |= key_bits[nxt]\n            next_state = next_mask * n + nxt\n            next_distance = distance + costs[nxt]\n\n            if next_distance < distances[next_state]:\n                distances[next_state] = next_distance\n                heappush(heap, (next_distance, next_state))\n\n    return -1\n```"}
{"model": "typesafe/jev-router", "task": "X1", "tier": "extreme", "rep": 1, "pass": false, "detail": "12,11,7", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.013678, "in_tok": 364, "out_tok": 1295, "latency": 25.65, "id": "gen-1791298176-Ym8ASoleQRBBKX5H3Hq0", "text": "<answer>12,11,7</answer>"}
{"model": "typesafe/jev-router", "task": "X3", "tier": "extreme", "rep": 2, "pass": true, "detail": "66/66 hidden tests", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.012662, "in_tok": 276, "out_tok": 1211, "latency": 25.31, "id": "gen-1791298184-Jok7IbeVL0d3hos7zGjy", "text": "```python\ndef vault2(grid: list[str]) -> int:\n    from heapq import heappop, heappush\n\n    if not grid or not grid[0]:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    size = rows * cols\n    cells = \"\".join(grid)\n\n    start = cells.find(\"S\")\n    end = cells.find(\"E\")\n    if start < 0 or end < 0:\n        return -1\n\n    # Each bit records either a collected key cell or an opened door cell.\n    key_bits = [0] * size\n    door_bits = [0] * size\n    letters = [-1] * size\n    key_masks = [0] * 3\n    door_masks = [0] * 3\n    bit_count = 0\n\n    for pos, cell in enumerate(cells):\n        if cell in \"abc\":\n            letter = ord(cell) - ord(\"a\")\n            bit = 1 << bit_count\n            bit_count += 1\n            key_bits[pos] = bit\n            key_masks[letter] |= bit\n            letters[pos] = letter\n        elif cell in \"ABC\":\n            letter = ord(cell) - ord(\"A\")\n            bit = 1 << bit_count\n            bit_count += 1\n            door_bits[pos] = bit\n            door_masks[letter] |= bit\n            letters[pos] = letter\n\n    neighbors = [[] for _ in range(size)]\n    for pos, cell in enumerate(cells):\n        if cell == \"#\":\n            continue\n        row, col = divmod(pos, cols)\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = row + dr, col + dc\n            if 0 <= nr < rows and 0 <= nc < cols:\n                nxt = nr * cols + nc\n                target = cells[nxt]\n                if target != \"#\":\n                    cost = 1 + (int(target) if target in \"123456789\" else 0)\n                    neighbors[pos].append((nxt, cost))\n\n    inf = float(\"inf\")\n    distances = [inf] * (size * (1 << bit_count))\n    distances[start] = 0\n    heap = [(0, start)]\n\n    while heap:\n        cost, state = heappop(heap)\n        if cost != distances[state]:\n            continue\n\n        mask, pos = divmod(state, size)\n        if pos == end:\n            return cost\n\n        for nxt, step_cost in neighbors[pos]:\n            next_mask = mask\n            door_bit = door_bits[nxt]\n\n            if door_bit and not (mask & door_bit):\n                letter = letters[nxt]\n                collected = (mask & key_masks[letter]).bit_count()\n                consumed = (mask & door_masks[letter]).bit_count()\n                if collected <= consumed:\n                    continue\n                next_mask |= door_bit\n\n            next_mask |= key_bits[nxt]\n            next_state = next_mask * size + nxt\n            next_cost = cost + step_cost\n\n            if next_cost < distances[next_state]:\n                distances[next_state] = next_cost\n                heappush(heap, (next_cost, next_state))\n\n    return -1\n```"}
{"model": "nvidia/switchyard", "task": "X3", "tier": "extreme", "rep": 1, "pass": true, "detail": "66/66 hidden tests", "routed": "moonshotai/kimi-k3", "provider": "Relace", "cost": 0.0286831, "in_tok": 355, "out_tok": 2184, "latency": 38.52, "id": "gen-1791298199-qfuaedBSBfkimIlaksI0", "text": "```python\nimport heapq\n\ndef vault2(grid: list[str]) -> int:\n    R, C = len(grid), len(grid[0])\n    start = end = None\n    key_idx = {}    # (r,c) -> index into keys\n    door_idx = {}   # (r,c) -> index into doors\n    keys = []       # letter of each key cell\n    doors = []      # lowercase letter of each door cell\n\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            if ch == 'S':\n                start = (r, c)\n            elif ch == 'E':\n                end = (r, c)\n            elif ch in 'abc':\n                key_idx[(r, c)] = len(keys)\n                keys.append(ch)\n            elif ch in 'ABC':\n                door_idx[(r, c)] = len(doors)\n                doors.append(ch.lower())\n\n    K, D = len(keys), len(doors)\n\n    def enter_cost(ch):\n        return 1 + int(ch) if ch.isdigit() else 1\n\n    def available(letter, km, dm):\n        got = sum(1 for j in range(K) if keys[j] == letter and (km >> j) & 1)\n        used = sum(1 for j in range(D) if doors[j] == letter and (dm >> j) & 1)\n        return got - used\n\n    sr, sc = start\n    dist = {(sr, sc, 0, 0): 0}\n    pq = [(0, sr, sc, 0, 0)]\n    DIRS = ((1, 0), (-1, 0), (0, 1), (0, -1))\n\n    while pq:\n        cost, r, c, km, dm = heapq.heappop(pq)\n        if cost > dist.get((r, c, km, dm), float('inf')):\n            continue\n        if (r, c) == end:\n            return cost\n        for dr, dc in DIRS:\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n            ch = grid[nr][nc]\n            if ch == '#':\n                continue\n            nkm, ndm = km, dm\n            di = door_idx.get((nr, nc))\n            if di is not None and not (dm >> di) & 1:\n                if available(doors[di], km, dm) < 1:\n                    continue\n                ndm = dm | (1 << di)\n            else:\n                ki = key_idx.get((nr, nc))\n                if ki is not None and not (km >> ki) & 1:\n                    nkm = km | (1 << ki)\n            ncost = cost + enter_cost(ch)\n            state = (nr, nc, nkm, ndm)\n            if ncost < dist.get(state, float('inf')):\n                dist[state] = ncost\n                heapq.heappush(pq, (ncost, nr, nc, nkm, ndm))\n\n    return -1\n```"}
{"model": "nvidia/switchyard", "task": "X2", "tier": "extreme", "rep": 1, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0436965, "in_tok": 279, "out_tok": 2848, "latency": 47.87, "id": "gen-1791298199-ktobciUu0c8cAkA5b5yZ", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "nvidia/switchyard", "task": "X3", "tier": "extreme", "rep": 0, "pass": true, "detail": "66/66 hidden tests", "routed": "moonshotai/kimi-k3", "provider": "Fireworks", "cost": 0.052794, "in_tok": 358, "out_tok": 3448, "latency": 77.41, "id": "gen-1791298188-dfV2FoRQ4i8Ac1pOS2YH", "text": "```python\nimport heapq\n\ndef vault2(grid: list[str]) -> int:\n    if not grid:\n        return -1\n    R, C = len(grid), len(grid[0])\n\n    start = end = None\n    key_index = {}    # (r, c) -> key bit index\n    door_index = {}   # (r, c) -> door bit index\n    key_letter = []   # per key index: 'a'/'b'/'c'\n    door_letter = []  # per door index: 'A'/'B'/'C'\n\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            if ch == 'S':\n                start = (r, c)\n            elif ch == 'E':\n                end = (r, c)\n            elif ch in 'abc':\n                key_index[(r, c)] = len(key_letter)\n                key_letter.append(ch)\n            elif ch in 'ABC':\n                door_index[(r, c)] = len(door_letter)\n                door_letter.append(ch)\n\n    if start is None or end is None:\n        return -1\n\n    # Cost of entering each cell (paid on every entry).\n    cost = [[0] * C for _ in range(R)]\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            cost[r][c] = 1 + (ord(ch) - 48) if '1' <= ch <= '9' else 1\n\n    # Per-letter bitmasks so available keys of letter L =\n    # popcount(collected & key_mask[L]) - popcount(opened & door_mask[L]).\n    key_mask = {'a': 0, 'b': 0, 'c': 0}\n    for i, ch in enumerate(key_letter):\n        key_mask[ch] |= 1 << i\n    door_mask = {'A': 0, 'B': 0, 'C': 0}\n    for j, ch in enumerate(door_letter):\n        door_mask[ch] |= 1 << j\n\n    sr, sc = start\n    er, ec = end\n    dist = {(sr, sc, 0, 0): 0}\n    pq = [(0, sr, sc, 0, 0)]  # (cost, r, c, collected_keys, opened_doors)\n    INF = float('inf')\n\n    while pq:\n        d, r, c, km, om = heapq.heappop(pq)\n        if d != dist.get((r, c, km, om)):\n            continue\n        if (r, c) == (er, ec):\n            return d\n        for dr, dc in ((1, 0), (-1, 0), (0, 1), (0, -1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n            if grid[nr][nc] == '#':\n                continue\n            nkm, nom = km, om\n            pos = (nr, nc)\n            di = door_index.get(pos)\n            if di is not None:\n                if not (om >> di) & 1:\n                    L = door_letter[di]\n                    available = ((km & key_mask[L.lower()]).bit_count()\n                                 - (om & door_mask[L]).bit_count())\n                    if available <= 0:\n                        continue\n                    nom = om | (1 << di)\n            else:\n                ki = key_index.get(pos)\n                if ki is not None:\n                    nkm = km | (1 << ki)\n            nd = d + cost[nr][nc]\n            state = (nr, nc, nkm, nom)\n            if nd < dist.get(state, INF):\n                dist[state] = nd\n                heapq.heappush(pq, (nd, nr, nc, nkm, nom))\n\n    return -1\n```"}
{"model": "nvidia/switchyard", "task": "X4", "tier": "extreme", "rep": 0, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0492975, "in_tok": 345, "out_tok": 3206, "latency": 73.36, "id": "gen-1791298196-0LnDlhEE20VVLc7P3idU", "text": "<answer>{\"J1\": 4, \"J2\": 0, \"J3\": 10, \"J4\": 1, \"J5\": 10, \"J6\": 2, \"J7\": 14, \"J8\": 14, \"J9\": 18}</answer>"}
{"model": "nvidia/switchyard", "task": "X4", "tier": "extreme", "rep": 1, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0704775, "in_tok": 345, "out_tok": 4618, "latency": 70.61, "id": "gen-1791298200-lVXg8tTJTzTaUY47BgTO", "text": "<answer>{\"J1\": 4, \"J2\": 0, \"J3\": 10, \"J4\": 1, \"J5\": 10, \"J6\": 2, \"J7\": 14, \"J8\": 14, \"J9\": 18}</answer>"}
{"model": "nvidia/switchyard", "task": "X2", "tier": "extreme", "rep": 2, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0613665, "in_tok": 279, "out_tok": 4026, "latency": 63.55, "id": "gen-1791298209-KMpIvUDLRAZAeMWX4RDV", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "nvidia/switchyard", "task": "X3", "tier": "extreme", "rep": 2, "pass": true, "detail": "66/66 hidden tests", "routed": "moonshotai/kimi-k3", "provider": "Relace", "cost": 0.0292941, "in_tok": 355, "out_tok": 2231, "latency": 44.75, "id": "gen-1791298238-DGZKVfoUFJ3zxSE7Z6W1", "text": "```python\nimport heapq\n\n\ndef vault2(grid: list[str]) -> int:\n    R, C = len(grid), len(grid[0])\n    start = end = None\n    key_id = {}    # (r, c) -> index into keys\n    door_id = {}   # (r, c) -> index into doors\n    keys = []      # letter of each key cell\n    doors = []     # letter of each door cell\n\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            if ch == 'S':\n                start = (r, c)\n            elif ch == 'E':\n                end = (r, c)\n            elif ch in 'abc':\n                key_id[(r, c)] = len(keys)\n                keys.append(ch)\n            elif ch in 'ABC':\n                door_id[(r, c)] = len(doors)\n                doors.append(ch)\n\n    # State: (row, col, collected-key-cells mask, opened-door-cells mask).\n    # Keys carried of a letter = collected of that letter minus doors of that\n    # letter already opened (each opening consumed exactly one matching key).\n    INF = float('inf')\n    sr, sc = start\n    dist = {(sr, sc, 0, 0): 0}\n    pq = [(0, sr, sc, 0, 0)]\n\n    while pq:\n        d, r, c, km, dm = heapq.heappop(pq)\n        if d != dist.get((r, c, km, dm), INF):\n            continue\n        if (r, c) == end:\n            return d\n        for dr, dc in ((1, 0), (-1, 0), (0, 1), (0, -1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n            ch = grid[nr][nc]\n            if ch == '#':\n                continue\n\n            cost = 1 + (ord(ch) - ord('0')) if ch.isdigit() else 1\n            nkm, ndm = km, dm\n\n            pos = (nr, nc)\n            if pos in key_id:\n                bit = 1 << key_id[pos]\n                if not (km & bit):\n                    nkm = km | bit          # pick up this key (first visit)\n            elif pos in door_id:\n                bit = 1 << door_id[pos]\n                if not (dm & bit):          # door not yet opened: needs a key\n                    letter = ch.lower()\n                    carried = sum(\n                        1 for i, k in enumerate(keys)\n                        if k == letter and (km >> i) & 1\n                    ) - sum(\n                        1 for i, dl in enumerate(doors)\n                        if dl.lower() == letter and (dm >> i) & 1\n                    )\n                    if carried < 1:\n                        continue\n                    ndm = dm | bit          # consume key, open door permanently\n\n            nd = d + cost\n            state = (nr, nc, nkm, ndm)\n            if nd < dist.get(state, INF):\n                dist[state] = nd\n                heapq.heappush(pq, (nd, nr, nc, nkm, ndm))\n\n    return -1\n```"}
{"model": "nvidia/switchyard", "task": "X2", "tier": "extreme", "rep": 0, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0862965, "in_tok": 279, "out_tok": 5688, "latency": 97.82, "id": "gen-1791298187-3f22RGc2Q8YRnQCWAqrH", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "unbiased/pareto", "task": "X2", "tier": "extreme", "rep": 0, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00269925, "in_tok": 213, "out_tok": 291, "latency": 41.3, "id": "gen-1791298269-PtxoeTeYX5oTXJTvwIlS", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "nvidia/switchyard", "task": "X1", "tier": "extreme", "rep": 0, "pass": true, "detail": "1,4,4,-132,11,7", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.1445455, "in_tok": 443, "out_tok": 9533, "latency": 125.66, "id": "gen-1791298186-nja0iol7stYHEzyh4zIN", "text": "<answer>1,4,4,-132,11,7</answer>"}
{"model": "unbiased/pareto", "task": "X2", "tier": "extreme", "rep": 1, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00259275, "in_tok": 213, "out_tok": 296, "latency": 29.34, "id": "gen-1791298285-DDfUUorHnOfbDKFqtPaI", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "nvidia/switchyard", "task": "X4", "tier": "extreme", "rep": 2, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.0975525, "in_tok": 345, "out_tok": 6423, "latency": 77.46, "id": "gen-1791298247-gZq6ZGoVGhyoOMJwdbXn", "text": "<answer>{\"J1\": 4, \"J2\": 0, \"J3\": 10, \"J4\": 6, \"J5\": 10, \"J6\": 1, \"J7\": 14, \"J8\": 14, \"J9\": 18}</answer>"}
{"model": "unbiased/pareto", "task": "X4", "tier": "extreme", "rep": 0, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00351925, "in_tok": 286, "out_tok": 376, "latency": 57.13, "id": "gen-1791298273-pW7Ln9Dt251QOhoSkFiA", "text": "<answer>{\"J1\":4,\"J2\":0,\"J3\":10,\"J4\":1,\"J5\":10,\"J6\":0,\"J7\":14,\"J8\":14,\"J9\":18}</answer>"}
{"model": "unbiased/pareto", "task": "X3", "tier": "extreme", "rep": 0, "pass": true, "detail": "66/66 hidden tests", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00812425, "in_tok": 295, "out_tok": 987, "latency": 59.32, "id": "gen-1791298271-CUpHaY8N4vt7f9RfH6g1", "text": "```python\ndef vault2(grid: list[str]) -> int:\n    from heapq import heappop, heappush\n\n    if not grid or not grid[0]:\n        return -1\n\n    height, width = len(grid), len(grid[0])\n    cells = \"\".join(grid)\n    size = len(cells)\n\n    start = cells.find(\"S\")\n    end = cells.find(\"E\")\n    if start == -1 or end == -1:\n        return -1\n\n    key_bit = [0] * size\n    door_bit = [0] * size\n    door_kind = [-1] * size\n    keys_by_kind = [0] * 3\n    doors_by_kind = [0] * 3\n    entry_cost = [1] * size\n    key_count = door_count = 0\n\n    for pos, cell in enumerate(cells):\n        if cell in \"abc\":\n            kind = ord(cell) - ord(\"a\")\n            bit = 1 << key_count\n            key_count += 1\n            key_bit[pos] = bit\n            keys_by_kind[kind] |= bit\n        elif cell in \"ABC\":\n            kind = ord(cell) - ord(\"A\")\n            bit = 1 << door_count\n            door_count += 1\n            door_bit[pos] = bit\n            door_kind[pos] = kind\n            doors_by_kind[kind] |= bit\n        elif cell in \"123456789\":\n            entry_cost[pos] += int(cell)\n\n    neighbors = [[] for _ in range(size)]\n    for pos, cell in enumerate(cells):\n        if cell == \"#\":\n            continue\n        row, col = divmod(pos, width)\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = row + dr, col + dc\n            if 0 <= nr < height and 0 <= nc < width:\n                nxt = nr * width + nc\n                if cells[nxt] != \"#\":\n                    neighbors[pos].append(nxt)\n\n    # Inventory is determined by collected key cells minus opened door cells.\n    best = {(start, 0, 0): 0}\n    heap = [(0, start, 0, 0)]\n\n    while heap:\n        cost, pos, picked, opened = heappop(heap)\n        if best.get((pos, picked, opened)) != cost:\n            continue\n        if pos == end:\n            return cost\n\n        for nxt in neighbors[pos]:\n            next_opened = opened\n            bit = door_bit[nxt]\n            if bit and not (opened & bit):\n                kind = door_kind[nxt]\n                collected = (picked & keys_by_kind[kind]).bit_count()\n                spent = (opened & doors_by_kind[kind]).bit_count()\n                if collected <= spent:\n                    continue\n                next_opened |= bit\n\n            next_picked = picked | key_bit[nxt]\n            next_cost = cost + entry_cost[nxt]\n            state = (nxt, next_picked, next_opened)\n            previous = best.get(state)\n\n            if previous is None or next_cost < previous:\n                best[state] = next_cost\n                heappush(heap, (next_cost, nxt, next_picked, next_opened))\n\n    return -1\n```"}
{"model": "nvidia/switchyard", "task": "X1", "tier": "extreme", "rep": 1, "pass": true, "detail": "1,4,4,-132,11,7", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.1815205, "in_tok": 443, "out_tok": 11998, "latency": 141.07, "id": "gen-1791298197-H0skhsJVkmIbsRhz8BYt", "text": "<answer>1,4,4,-132,11,7</answer>"}
{"model": "unbiased/pareto", "task": "X1", "tier": "extreme", "rep": 0, "pass": true, "detail": "1,4,4,-132,11,7", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00917725, "in_tok": 382, "out_tok": 1156, "latency": 79.58, "id": "gen-1791298266-P1m5ZRYfV3XvtvPsyE1r", "text": "<answer>1,4,4,-132,11,7</answer>"}
{"model": "unbiased/pareto", "task": "X2", "tier": "extreme", "rep": 2, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00262275, "in_tok": 213, "out_tok": 300, "latency": 28.15, "id": "gen-1791298324-ATYWoTni79Vh7rvHMmn5", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "openrouter/auto", "task": "X2", "tier": "extreme", "rep": 0, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0030222, "in_tok": 226, "out_tok": 2462, "latency": 10.81, "id": "gen-1791298345-9zctRGNnZxf8oh68HCXk", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "openrouter/auto", "task": "X4", "tier": "extreme", "rep": 0, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0030321, "in_tok": 291, "out_tok": 2454, "latency": 9.57, "id": "gen-1791298356-ACUsFCC9TTZKS8Dl3Y9e", "text": "<answer>{\"J1\": 4, \"J2\": 0, \"J3\": 10, \"J4\": 6, \"J5\": 10, \"J6\": 1, \"J7\": 14, \"J8\": 14, \"J9\": 18}</answer>"}
{"model": "unbiased/pareto", "task": "X3", "tier": "extreme", "rep": 1, "pass": true, "detail": "66/66 hidden tests", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00698725, "in_tok": 295, "out_tok": 917, "latency": 56.76, "id": "gen-1791298311-hwzQJHVyFuYTHoIygedk", "text": "```python\ndef vault2(grid: list[str]) -> int:\n    from heapq import heappop, heappush\n\n    if not grid or not grid[0]:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    size = rows * cols\n    cells = \"\".join(grid)\n\n    start = cells.find(\"S\")\n    end = cells.find(\"E\")\n    if start == -1 or end == -1:\n        return -1\n\n    key_bits = [0] * size\n    door_bits = [0] * size\n    door_types = [-1] * size\n    keys_by_type = [0] * 3\n    doors_by_type = [0] * 3\n    costs = [1] * size\n\n    key_count = door_count = 0\n    for pos, cell in enumerate(cells):\n        if cell in \"123456789\":\n            costs[pos] += int(cell)\n        elif cell in \"abc\":\n            letter = ord(cell) - ord(\"a\")\n            bit = 1 << key_count\n            key_count += 1\n            key_bits[pos] = bit\n            keys_by_type[letter] |= bit\n        elif cell in \"ABC\":\n            letter = ord(cell) - ord(\"A\")\n            bit = 1 << door_count\n            door_count += 1\n            door_bits[pos] = bit\n            door_types[pos] = letter\n            doors_by_type[letter] |= bit\n\n    neighbors = [[] for _ in range(size)]\n    for pos, cell in enumerate(cells):\n        if cell == \"#\":\n            continue\n        row, col = divmod(pos, cols)\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = row + dr, col + dc\n            if 0 <= nr < rows and 0 <= nc < cols:\n                nxt = nr * cols + nc\n                if cells[nxt] != \"#\":\n                    neighbors[pos].append(nxt)\n\n    # Inventory is determined by collected key cells minus opened door cells.\n    distances = {(start, 0, 0): 0}\n    heap = [(0, start, 0, 0)]\n\n    while heap:\n        cost, pos, collected, opened = heappop(heap)\n        if distances.get((pos, collected, opened)) != cost:\n            continue\n        if pos == end:\n            return cost\n\n        for nxt in neighbors[pos]:\n            next_collected = collected | key_bits[nxt]\n            next_opened = opened\n\n            door_bit = door_bits[nxt]\n            if door_bit and not (opened & door_bit):\n                letter = door_types[nxt]\n                acquired = (collected & keys_by_type[letter]).bit_count()\n                consumed = (opened & doors_by_type[letter]).bit_count()\n                if acquired <= consumed:\n                    continue\n                next_opened |= door_bit\n\n            next_cost = cost + costs[nxt]\n            state = (nxt, next_collected, next_opened)\n            previous = distances.get(state)\n            if previous is None or next_cost < previous:\n                distances[state] = next_cost\n                heappush(\n                    heap, (next_cost, nxt, next_collected, next_opened)\n                )\n\n    return -1\n```"}
{"model": "unbiased/pareto", "task": "X4", "tier": "extreme", "rep": 1, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00264475, "in_tok": 286, "out_tok": 341, "latency": 59.04, "id": "gen-1791298312-aeK5OmAr4wGEPIqX7jSh", "text": "<answer>{\"J1\":4,\"J2\":0,\"J3\":10,\"J4\":1,\"J5\":10,\"J6\":0,\"J7\":14,\"J8\":14,\"J9\":18}</answer>"}
{"model": "unbiased/pareto", "task": "X1", "tier": "extreme", "rep": 1, "pass": true, "detail": "1,4,4,-132,11,7", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00803875, "in_tok": 382, "out_tok": 1057, "latency": 88.65, "id": "gen-1791298283-U1m70Ym0fB0U9u7jokvf", "text": "<answer>1,4,4,-132,11,7</answer>"}
{"model": "unbiased/pareto", "task": "X4", "tier": "extreme", "rep": 2, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00229225, "in_tok": 286, "out_tok": 294, "latency": 46.56, "id": "gen-1791298330-pRKpYzhjU396ApK26zKM", "text": "<answer>{\"J1\":4,\"J2\":0,\"J3\":10,\"J4\":1,\"J5\":10,\"J6\":0,\"J7\":14,\"J8\":14,\"J9\":18}</answer>"}
{"model": "openrouter/auto", "task": "X2", "tier": "extreme", "rep": 1, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0035454, "in_tok": 226, "out_tok": 2898, "latency": 11.69, "id": "gen-1791298368-BWQD2rbJRzDurHp7TeM9", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "openrouter/auto", "task": "X4", "tier": "extreme", "rep": 1, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.002696436, "in_tok": 291, "out_tok": 2237, "latency": 8.56, "id": "gen-1791298371-VFH4w4gKwsIbHwsjAD5n", "text": "<answer>{\"J1\":4,\"J2\":0,\"J3\":10,\"J4\":1,\"J5\":10,\"J6\":2,\"J7\":14,\"J8\":14,\"J9\":18}</answer>"}
{"model": "openrouter/auto", "task": "X3", "tier": "extreme", "rep": 0, "pass": true, "detail": "66/66 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0090186, "in_tok": 306, "out_tok": 7439, "latency": 29.02, "id": "gen-1791298352-mhN2wzW8RsUlC6KGgJmx", "text": "```python\nimport heapq\nfrom collections import defaultdict\n\ndef vault2(grid: list[str]) -> int:\n    if not grid or not grid[0]:\n        return -1\n\n    R, C = len(grid), len(grid[0])\n    sr = sc = er = ec = None\n\n    keys = []\n    doors = []\n    key_at = {}\n    door_at = {}\n\n    for r in range(R):\n        for c, ch in enumerate(grid[r]):\n            if ch == 'S':\n                sr, sc = r, c\n            elif ch == 'E':\n                er, ec = r, c\n            elif ch.islower():\n                key_at[(r, c)] = len(keys)\n                keys.append((r, c, ch))\n            elif ch.isupper() and ch not in 'SE':\n                door_at[(r, c)] = len(doors)\n                doors.append((r, c, ch.lower()))\n\n    key_by_letter = defaultdict(list)\n    for i, (_, _, ch) in enumerate(keys):\n        key_by_letter[ch].append(i)\n\n    door_by_letter = defaultdict(list)\n    for j, (_, _, ch) in enumerate(doors):\n        door_by_letter[ch].append(j)\n\n    INF = 10**18\n    start = (sr, sc, 0, 0)\n    dist = {start: 0}\n    pq = [(0, sr, sc, 0, 0)]\n\n    directions = [(1, 0), (-1, 0), (0, 1), (0, -1)]\n\n    while pq:\n        d, r, c, picked, opened = heapq.heappop(pq)\n        state = (r, c, picked, opened)\n\n        if d != dist.get(state):\n            continue\n\n        if (r, c) == (er, ec):\n            return d\n\n        for dr, dc in directions:\n            nr, nc = r + dr, c + dc\n\n            if nr < 0 or nr >= R or nc < 0 or nc >= C:\n                continue\n\n            ch = grid[nr][nc]\n            if ch == '#':\n                continue\n\n            npicked = picked\n            nopened = opened\n            add = 1\n\n            if ch.isdigit():\n                add = 1 + int(ch)\n\n            elif ch.islower():\n                i = key_at[(nr, nc)]\n                if not ((picked >> i) & 1):\n                    npicked = picked | (1 << i)\n\n            elif ch.isupper() and ch not in 'SE':\n                j = door_at[(nr, nc)]\n\n                if (opened >> j) & 1:\n                    add = 0\n                else:\n                    letter = ch.lower()\n                    have = 0\n\n                    for i in key_by_letter.get(letter, ()):\n                        if (picked >> i) & 1:\n                            have += 1\n\n                    for jj in door_by_letter.get(letter, ()):\n                        if (opened >> jj) & 1:\n                            have -= 1\n\n                    if have <= 0:\n                        continue\n\n                    nopened = opened | (1 << j)\n                    add = 1\n\n            nd = d + add\n            nstate = (nr, nc, npicked, nopened)\n\n            if nd < dist.get(nstate, INF):\n                dist[nstate] = nd\n                heapq.heappush(pq, (nd, nr, nc, npicked, nopened))\n\n    return -1\n```"}
{"model": "openrouter/auto", "task": "X1", "tier": "extreme", "rep": 0, "pass": true, "detail": "1,4,4,-132,11,7", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0185721, "in_tok": 391, "out_tok": 15379, "latency": 45.82, "id": "gen-1791298338-gorqLA1FxPruarD7EtcA", "text": "<answer>1,4,4,-132,11,7</answer>"}
{"model": "openrouter/auto", "task": "X2", "tier": "extreme", "rep": 2, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.0026682, "in_tok": 226, "out_tok": 2167, "latency": 8.61, "id": "gen-1791298379-rxLNS9Cddp1Ao2ONiPvX", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "openrouter/auto", "task": "X4", "tier": "extreme", "rep": 2, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.002574036, "in_tok": 291, "out_tok": 2135, "latency": 8.5, "id": "gen-1791298381-RB4c6uTNMgov3D0F1Xyf", "text": "<answer>{\"J1\": 4, \"J2\": 0, \"J3\": 10, \"J4\": 1, \"J5\": 10, \"J6\": 2, \"J7\": 14, \"J8\": 14, \"J9\": 18}</answer>"}
{"model": "unbiased/pareto", "task": "X3", "tier": "extreme", "rep": 2, "pass": true, "detail": "66/66 hidden tests", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.00844675, "in_tok": 295, "out_tok": 1030, "latency": 63.59, "id": "gen-1791298330-no2rJeZSxC3qTo2Rwpe8", "text": "```python\ndef vault2(grid: list[str]) -> int:\n    from heapq import heappop, heappush\n\n    if not grid or not grid[0]:\n        return -1\n\n    rows, cols = len(grid), len(grid[0])\n    cells = \"\".join(grid)\n    start, end = cells.find(\"S\"), cells.find(\"E\")\n    if start == -1 or end == -1:\n        return -1\n\n    n = len(cells)\n    key_bit = [0] * n\n    door_bit = [0] * n\n    door_type = [-1] * n\n    key_masks = [0] * 3\n    door_masks = [0] * 3\n    entry_cost = [1] * n\n    key_count = door_count = 0\n\n    for pos, cell in enumerate(cells):\n        if cell in \"abc\":\n            kind = ord(cell) - ord(\"a\")\n            bit = 1 << key_count\n            key_count += 1\n            key_bit[pos] = bit\n            key_masks[kind] |= bit\n        elif cell in \"ABC\":\n            kind = ord(cell) - ord(\"A\")\n            bit = 1 << door_count\n            door_count += 1\n            door_bit[pos] = bit\n            door_type[pos] = kind\n            door_masks[kind] |= bit\n        elif cell in \"123456789\":\n            entry_cost[pos] += int(cell)\n\n    neighbors = [[] for _ in range(n)]\n    for pos, cell in enumerate(cells):\n        if cell == \"#\":\n            continue\n        r, c = divmod(pos, cols)\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = r + dr, c + dc\n            if 0 <= nr < rows and 0 <= nc < cols:\n                nxt = nr * cols + nc\n                if cells[nxt] != \"#\":\n                    neighbors[pos].append(nxt)\n\n    # Inventory is determined by collected key cells minus opened door cells.\n    collected_counts = [\n        tuple((mask & typed_mask).bit_count() for typed_mask in key_masks)\n        for mask in range(1 << key_count)\n    ]\n    spent_counts = [\n        tuple((mask & typed_mask).bit_count() for typed_mask in door_masks)\n        for mask in range(1 << door_count)\n    ]\n\n    distances = {(start, 0, 0): 0}\n    heap = [(0, start, 0, 0)]\n\n    while heap:\n        cost, pos, collected, opened = heappop(heap)\n        if cost != distances.get((pos, collected, opened)):\n            continue\n        if pos == end:\n            return cost\n\n        for nxt in neighbors[pos]:\n            next_opened = opened\n            bit = door_bit[nxt]\n            if bit and not (opened & bit):\n                kind = door_type[nxt]\n                if collected_counts[collected][kind] <= spent_counts[opened][kind]:\n                    continue\n                next_opened |= bit\n\n            next_collected = collected | key_bit[nxt]\n            next_cost = cost + entry_cost[nxt]\n            state = (nxt, next_collected, next_opened)\n            previous = distances.get(state)\n            if previous is None or next_cost < previous:\n                distances[state] = next_cost\n                heappush(heap, (next_cost, nxt, next_collected, next_opened))\n\n    return -1\n```"}
{"model": "unbiased/pareto", "task": "X1", "tier": "extreme", "rep": 2, "pass": true, "detail": "1,4,4,-132,11,7", "routed": "unbiased/pareto", "provider": "Unbiased", "cost": 0.01098625, "in_tok": 382, "out_tok": 1450, "latency": 86.03, "id": "gen-1791298314-rOab62dPuLpHOJA5eEgd", "text": "<answer>1,4,4,-132,11,7</answer>"}
{"model": "openrouter/auto", "task": "X3", "tier": "extreme", "rep": 2, "pass": false, "detail": "42/66 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.008481336, "in_tok": 306, "out_tok": 7054, "latency": 23.64, "id": "gen-1791298380-uqevMIdBpAetiUbYukiR", "text": "```python\ndef vault2(grid: list[str]) -> int:\n    import heapq\n\n    if not grid:\n        return -1\n\n    R, C = len(grid), len(grid[0])\n    N = R * C\n\n    walls = [False] * N\n    digit = [0] * N\n    key_bit = [-1] * N\n    door_bit = [-1] * N\n\n    key_letter = []\n    door_letter = []\n\n    start = end = -1\n\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            idx = r * C + c\n\n            if ch == '#':\n                walls[idx] = True\n            elif ch == 'S':\n                start = idx\n            elif ch == 'E':\n                end = idx\n            elif '1' <= ch <= '9':\n                digit[idx] = ord(ch) - ord('0')\n            elif 'a' <= ch <= 'c':\n                key_bit[idx] = len(key_letter)\n                key_letter.append(ch)\n            elif 'A' <= ch <= 'C':\n                door_bit[idx] = len(door_letter)\n                door_letter.append(ch)\n\n    K = len(key_letter)\n    D = len(door_letter)\n\n    INF = 10**9\n    dist = [[[INF] * (1 << D) for _ in range(1 << K)] for _ in range(N)]\n    dist[start][0][0] = 0\n\n    heap = [(0, start, 0, 0)]\n    dirs = ((1, 0), (-1, 0), (0, 1), (0, -1))\n\n    while heap:\n        cur, pos, km, dm = heapq.heappop(heap)\n\n        if cur != dist[pos][km][dm]:\n            continue\n\n        if pos == end:\n            return cur\n\n        r, c = divmod(pos, C)\n\n        for dr, dc in dirs:\n            nr, nc = r + dr, c + dc\n\n            if nr < 0 or nr >= R or nc < 0 or nc >= C:\n                continue\n\n            np = nr * C + nc\n\n            if walls[np]:\n                continue\n\n            nkm = km\n            ndm = dm\n\n            if door_bit[np] != -1:\n                db = door_bit[np]\n\n                if (dm & (1 << db)) == 0:\n                    letter = door_letter[db]\n\n                    collected = 0\n                    for i, l in enumerate(key_letter):\n                        if (km & (1 << i)) and l == letter:\n                            collected += 1\n\n                    opened = 0\n                    for i, l in enumerate(door_letter):\n                        if (dm & (1 << i)) and l == letter:\n                            opened += 1\n\n                    if collected <= opened:\n                        continue\n\n                    ndm = dm | (1 << db)\n\n            if key_bit[np] != -1:\n                kb = key_bit[np]\n                if (km & (1 << kb)) == 0:\n                    nkm = km | (1 << kb)\n\n            cost = 1 + digit[np]\n            nd = cur + cost\n\n            if nd < dist[np][nkm][ndm]:\n                dist[np][nkm][ndm] = nd\n                heapq.heappush(heap, (nd, np, nkm, ndm))\n\n    return -1\n```"}
{"model": "openrouter/auto", "task": "X3", "tier": "extreme", "rep": 1, "pass": false, "detail": "30/66 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.012170136, "in_tok": 306, "out_tok": 10128, "latency": 36.5, "id": "gen-1791298371-F10MRQlybCRfH20guC5t", "text": "```python\nfrom heapq import heappush, heappop\n\ndef vault2(grid: list[str]) -> int:\n    if not grid or not grid[0]:\n        return -1\n\n    H, W = len(grid), len(grid[0])\n\n    key_id = [[-1] * W for _ in range(H)]\n    door_id = [[-1] * W for _ in range(H)]\n    key_letter_masks = {}\n    door_letter_masks = {}\n\n    sr = sc = er = ec = -1\n    K = D = 0\n\n    for r in range(H):\n        for c, ch in enumerate(grid[r]):\n            if ch == 'S':\n                sr, sc = r, c\n            elif ch == 'E':\n                er, ec = r, c\n            elif 'a' <= ch <= 'z':\n                key_id[r][c] = K\n                key_letter_masks[ch] = key_letter_masks.get(ch, 0) | (1 << K)\n                K += 1\n            elif 'A' <= ch <= 'Z':\n                door_id[r][c] = D\n                door_letter_masks[ch] = door_letter_masks.get(ch, 0) | (1 << D)\n                D += 1\n\n    if sr == er and sc == ec:\n        return 0\n\n    KM = 1 << K\n    DM = 1 << D\n    INF = 10**18\n\n    def state_index(r: int, c: int, km: int, dm: int) -> int:\n        return (((r * W + c) * KM + km) * DM + dm)\n\n    dist = [INF] * (H * W * KM * DM)\n    start = state_index(sr, sc, 0, 0)\n    dist[start] = 0\n\n    heap = [(0, sr, sc, 0, 0)]\n    directions = [(1, 0), (-1, 0), (0, 1), (0, -1)]\n\n    while heap:\n        cost, r, c, km, dm = heappop(heap)\n        cur = state_index(r, c, km, dm)\n\n        if cost != dist[cur]:\n            continue\n\n        if grid[r][c] == 'E':\n            return cost\n\n        for dr, dc in directions:\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < H and 0 <= nc < W):\n                continue\n\n            ch = grid[nr][nc]\n            if ch == '#':\n                continue\n\n            add = 1\n            if '1' <= ch <= '9':\n                add += ord(ch) - ord('0')\n\n            nkm = km\n            ndm = dm\n\n            if 'a' <= ch <= 'z':\n                kid = key_id[nr][nc]\n                if kid != -1:\n                    bit = 1 << kid\n                    if not (nkm & bit):\n                        nkm |= bit\n\n            elif 'A' <= ch <= 'Z':\n                did = door_id[nr][nc]\n                if did != -1:\n                    bit = 1 << did\n                    if not (ndm & bit):\n                        letter = ch.lower()\n                        kmask = key_letter_masks.get(letter, 0)\n                        dmask = door_letter_masks.get(letter, 0)\n\n                        collected = (km & kmask).bit_count()\n                        opened = (dm & dmask).bit_count()\n\n                        if collected - opened <= 0:\n                            continue\n\n                        ndm |= bit\n\n            ni = state_index(nr, nc, nkm, ndm)\n            ncost = cost + add\n\n            if ncost < dist[ni]:\n                dist[ni] = ncost\n                heappush(heap, (ncost, nr, nc, nkm, ndm))\n\n    return -1\n```"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "X2", "tier": "extreme", "rep": 0, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "deepseek/deepseek-v4.1-flash", "provider": "InferenceNet", "cost": 0.00111922, "in_tok": 226, "out_tok": 3678, "latency": 22.53, "id": "gen-1791298388-9hT58aedZ1FL9G09UikS", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "nvidia/switchyard", "task": "X1", "tier": "extreme", "rep": 2, "pass": true, "detail": "1,4,4,-132,11,7", "routed": "moonshotai/kimi-k3", "provider": "InferenceNet", "cost": 0.2304205, "in_tok": 443, "out_tok": 15258, "latency": 209.92, "id": "gen-1791298201-RWu22d3jebD52SyrEF7H", "text": "<answer>1,4,4,-132,11,7</answer>"}
{"model": "openrouter/auto", "task": "X1", "tier": "extreme", "rep": 1, "pass": true, "detail": "1,4,4,-132,11,7", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.019823604, "in_tok": 391, "out_tok": 16516, "latency": 46.69, "id": "gen-1791298366-Yb0yoROqwl2yAU4S1Zt8", "text": "<answer>1,4,4,-132,11,7</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "X4", "tier": "extreme", "rep": 0, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 0.002400243, "in_tok": 291, "out_tok": 4183, "latency": 22.79, "id": "gen-1791298394-iKxzXzgLsNsBfIpPHu9U", "text": "<answer>{\"J1\": 4, \"J2\": 0, \"J3\": 10, \"J4\": 1, \"J5\": 10, \"J6\": 2, \"J7\": 14, \"J8\": 14, \"J9\": 18}</answer>"}
{"model": "google/gemini-3.8-flash", "task": "X2", "tier": "extreme", "rep": 0, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.00809475, "in_tok": 213, "out_tok": 2116, "latency": 16.3, "id": "gen-1791298404-HKLZtg2Fi2UD1TAgAesg", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "X2", "tier": "extreme", "rep": 0, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.003472, "in_tok": 201, "out_tok": 307, "latency": 8.97, "id": "gen-1791298412-YzGVvdZKSGtBk5rg5Eiv", "text": "<answer>Wick 5 Gale 1428</answer>"}
{"model": "google/gemini-3.8-flash", "task": "X3", "tier": "extreme", "rep": 0, "pass": true, "detail": "66/66 hidden tests", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.011163, "in_tok": 284, "out_tok": 2920, "latency": 18.37, "id": "gen-1791298407-NJ5l9GYMN1kTHHVLnYxH", "text": "```python\nimport heapq\n\n\ndef vault2(grid: list[str]) -> int:\n    R, C = len(grid), len(grid[0])\n\n    start = None\n    end = None\n    key_cells = {}\n    door_cells = {}\n\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            if ch == \"S\":\n                start = (r, c)\n            elif ch == \"E\":\n                end = (r, c)\n            elif ch.islower():\n                key_cells[(r, c)] = len(key_cells)\n            elif ch.isupper() and ch not in (\"S\", \"E\"):\n                door_cells[(r, c)] = len(door_cells)\n\n    if start == end:\n        return 0\n\n    sr, sc = start\n    dist = {(sr, sc, 0, 0): 0}\n    pq = [(0, sr, sc, 0, 0)]\n\n    while pq:\n        d, r, c, k_mask, d_mask = heapq.heappop(pq)\n\n        if (r, c) == end:\n            return d\n\n        if d > dist.get((r, c, k_mask, d_mask), float(\"inf\")):\n            continue\n\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n\n            cell = grid[nr][nc]\n            if cell == \"#\":\n                continue\n\n            cost = 1 + int(cell) if cell.isdigit() else 1\n            nk_mask = k_mask\n            nd_mask = d_mask\n\n            if (nr, nc) in key_cells:\n                nk_mask |= 1 << key_cells[(nr, nc)]\n\n            if (nr, nc) in door_cells:\n                d_idx = door_cells[(nr, nc)]\n                if not (d_mask & (1 << d_idx)):\n                    key_char = cell.lower()\n                    keys_collected = sum(\n                        1\n                        for pos, kidx in key_cells.items()\n                        if grid[pos[0]][pos[1]] == key_char\n                        and (nk_mask & (1 << kidx))\n                    )\n                    doors_opened = sum(\n                        1\n                        for pos, didx in door_cells.items()\n                        if grid[pos[0]][pos[1]] == cell and (d_mask & (1 << didx))\n                    )\n                    if keys_collected - doors_opened <= 0:\n                        continue\n                    nd_mask |= 1 << d_idx\n\n            next_state = (nr, nc, nk_mask, nd_mask)\n            next_cost = d + cost\n            if next_cost < dist.get(next_state, float(\"inf\")):\n                dist[next_state] = next_cost\n                heapq.heappush(pq, (next_cost, nr, nc, nk_mask, nd_mask))\n\n    return -1\n```"}
{"model": "google/gemini-3.8-flash", "task": "X4", "tier": "extreme", "rep": 0, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.00920325, "in_tok": 266, "out_tok": 2401, "latency": 18.57, "id": "gen-1791298410-DtVBoYy4I57gFJtVQSDK", "text": "<answer>{\"J1\": 4, \"J2\": 0, \"J3\": 10, \"J4\": 5, \"J5\": 10, \"J6\": 0, \"J7\": 14, \"J8\": 14, \"J9\": 18}</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "X4", "tier": "extreme", "rep": 0, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.003784, "in_tok": 267, "out_tok": 325, "latency": 10.94, "id": "gen-1791298420-emn8zSCQFCOtDZ2WEMg1", "text": "<answer>{\"J1\":4,\"J2\":0,\"J3\":10,\"J4\":1,\"J5\":10,\"J6\":0,\"J7\":14,\"J8\":14,\"J9\":18}</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "X1", "tier": "extreme", "rep": 0, "pass": true, "detail": "1, 4, 4, -132, 11, 7", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.012308, "in_tok": 364, "out_tok": 1158, "latency": 22.8, "id": "gen-1791298411-SZO0eklKLmoDSwTDZvpx", "text": "<answer>1, 4, 4, -132, 11, 7</answer>"}
{"model": "openai/gpt-6.1-sol", "task": "X3", "tier": "extreme", "rep": 0, "pass": true, "detail": "66/66 hidden tests", "routed": "openai/gpt-6.1-sol", "provider": "OpenAI", "cost": 0.010932, "in_tok": 276, "out_tok": 1038, "latency": 19.41, "id": "gen-1791298417-WpQpg4WV9zeD1dMp9DAc", "text": "```python\ndef vault2(grid: list[str]) -> int:\n    import heapq\n\n    if not grid or not grid[0]:\n        return -1\n\n    height, width = len(grid), len(grid[0])\n    cells = \"\".join(grid)\n    size = len(cells)\n    start, end = cells.find(\"S\"), cells.find(\"E\")\n    if start < 0 or end < 0:\n        return -1\n\n    key_bits = [0] * size\n    door_bits = [0] * size\n    door_letters = [-1] * size\n    keys_by_letter = [0] * 3\n    doors_by_letter = [0] * 3\n    entry_cost = [1] * size\n\n    key_count = door_count = 0\n    for pos, cell in enumerate(cells):\n        if cell in \"abc\":\n            letter = ord(cell) - ord(\"a\")\n            bit = 1 << key_count\n            key_count += 1\n            key_bits[pos] = bit\n            keys_by_letter[letter] |= bit\n        elif cell in \"ABC\":\n            letter = ord(cell) - ord(\"A\")\n            bit = 1 << door_count\n            door_count += 1\n            door_bits[pos] = bit\n            door_letters[pos] = letter\n            doors_by_letter[letter] |= bit\n        elif \"1\" <= cell <= \"9\":\n            entry_cost[pos] += int(cell)\n\n    neighbors = [[] for _ in range(size)]\n    for pos in range(size):\n        if cells[pos] == \"#\":\n            continue\n        row, col = divmod(pos, width)\n        for dr, dc in ((-1, 0), (1, 0), (0, -1), (0, 1)):\n            nr, nc = row + dr, col + dc\n            if 0 <= nr < height and 0 <= nc < width:\n                nxt = nr * width + nc\n                if cells[nxt] != \"#\":\n                    neighbors[pos].append(nxt)\n\n    # Inventory is determined by collected key cells minus opened door cells.\n    initial = (start, 0, 0)\n    distances = {initial: 0}\n    heap = [(0, start, 0, 0)]\n\n    while heap:\n        cost, pos, collected, opened = heapq.heappop(heap)\n        state = (pos, collected, opened)\n        if cost != distances.get(state):\n            continue\n        if pos == end:\n            return cost\n\n        for nxt in neighbors[pos]:\n            next_opened = opened\n            door_bit = door_bits[nxt]\n            if door_bit and not (opened & door_bit):\n                letter = door_letters[nxt]\n                acquired = (collected & keys_by_letter[letter]).bit_count()\n                consumed = (opened & doors_by_letter[letter]).bit_count()\n                if acquired <= consumed:\n                    continue\n                next_opened |= door_bit\n\n            next_collected = collected | key_bits[nxt]\n            next_state = (nxt, next_collected, next_opened)\n            next_cost = cost + entry_cost[nxt]\n            previous = distances.get(next_state)\n            if previous is None or next_cost < previous:\n                distances[next_state] = next_cost\n                heapq.heappush(\n                    heap, (next_cost, nxt, next_collected, next_opened)\n                )\n\n    return -1\n```"}
{"model": "anthropic/claude-opus-5.5", "task": "X2", "tier": "extreme", "rep": 0, "pass": true, "detail": "Wick 5 Gale 1428", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.026244, "in_tok": 291, "out_tok": 1254, "latency": 11.94, "id": "gen-1791298426-b0hx95ynLDUwVM8M4t0B", "text": "Weekday: since 42453 = 9 \u00d7 4717, the weekday is unchanged (Wick).\n\n**Date**\n\n1. Year lengths are 31+28+30+33+29+31+30 = 212 days, or 214 in a leap year.\n2. From 22 Gale 1228 it is 8 days to 30 Gale, so 1 Ashen 1229 is 9 days later. That leaves 42453 \u2212 9 = 42444 days to count from 1 Ashen 1229.\n3. Any 40 consecutive years contain 8 multiples of 5 and exactly 1 multiple of 40, so 7 leap years. A 40-year block is therefore 40 \u00d7 212 + 7 \u00d7 2 = 8494 days.\n4. Four blocks take 4 \u00d7 8494 = 33976 days, reaching 1 Ashen 1389. The remainder is 42444 \u2212 33976 = 8468.\n5. For the years 1389\u20131427 (39 years):\n   - The multiples of 5 are 1390 to 1425, which is 8 years.\n   - Of these, 1400 is divisible by 40, so 7 are leap years.\n   - The total is 39 \u00d7 212 + 14 = 8282 days, reaching 1 Ashen 1428.\n   - The remainder is 8468 \u2212 8282 = 186.\n6. 1428 is not divisible by 5, so it is not a leap year. Counting 186 days from 1 Ashen:\n   - After Ashen (31): 155\n   - After Brine (28): 127\n   - After Cinder (30): 97\n   - After Dross (33): 64\n   - After Ember (29): 35\n   - After Fallow (31): 4\n   - 4 days past 1 Gale gives 5 Gale.\n\n<answer>Wick 5 Gale 1428</answer>"}
{"model": "openrouter/auto", "task": "X1", "tier": "extreme", "rep": 2, "pass": true, "detail": "1,4,4,-132,11,7", "routed": "deepseek/deepseek-v4.1-flash", "provider": "Together", "cost": 0.026488404, "in_tok": 391, "out_tok": 22070, "latency": 64.68, "id": "gen-1791298377-p6oTTnSlBN1L5LeUe6x4", "text": "<answer>1,4,4,-132,11,7</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "X1", "tier": "extreme", "rep": 0, "pass": true, "detail": "1,4,4,-132,11,7", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.036072, "in_tok": 518, "out_tok": 1700, "latency": 21.28, "id": "gen-1791298421-FWqcgUYN5QnZIcMJkBZD", "text": "<answer>1,4,4,-132,11,7</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "X4", "tier": "extreme", "rep": 0, "pass": true, "detail": "makespan 21 (optimal 21)", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.025272, "in_tok": 343, "out_tok": 1195, "latency": 11.92, "id": "gen-1791298431-TzVInqCI4PlOOemFlp16", "text": "The minimum makespan is 21. The proof has two parts: a lower bound, then a schedule that meets it.\n\n**Lower bound of 21**\n\n- J1 and J5 are both hot, so they cannot overlap.\n- J8 must wait for both J1 and J5 to finish.\n- J1 is released at 4 and lasts 6. J5 is released at 6 and lasts 4. They cannot both finish by 10 without overlapping, so one must run after the other.\n- **Option A:** J1 runs 4\u201310, then J5 runs 10\u201314. J8 then runs 14\u201318 and J9 runs 18\u201321. The chain ends at 21.\n- **Option B:** J5 runs 6\u201310, then J1 runs 10\u201316. J8 then runs 16\u201320 and J9 runs 20\u201323. The chain ends at 23.\n- So the makespan is at least 21.\n\n**Schedule achieving 21**\n\n| Machine | Job | Start | End |\n|---|---|---|---|\n| M1 | J1 | 4 | 10 |\n| M1 | J5 | 10 | 14 |\n| M1 | J8 | 14 | 18 |\n| M1 | J9 | 18 | 21 |\n| M2 | J2 | 0 | 1 |\n| M2 | J4 | 1 | 2 |\n| M2 | J6 | 2 | 7 |\n| M2 | J3 | 10 | 12 |\n| M2 | J7 | 14 | 16 |\n\n**Constraint checks**\n\n- **Release times:** J1 starts at 4, J3 at 10, J5 at 10 and J8 at 14. All are at or after their release times.\n- **Precedence:**\n  - J3 starts at 10, after J1 (ends 10) and J2 (ends 1).\n  - J8 starts at 14, after J1 (ends 10) and J5 (ends 14).\n  - J7 starts at 14, after J4 (ends 2), J5 (ends 14) and J6 (ends 7).\n  - J9 starts at 18, when J8 ends.\n- **Heat rule:** J2 runs 0\u20131, J1 runs 4\u201310 and J5 runs 10\u201314. None of the hot jobs overlap.\n- **Machines:** Each machine runs one job at a time.\n\n<answer>{\"J1\": 4, \"J2\": 0, \"J3\": 10, \"J4\": 1, \"J5\": 10, \"J6\": 2, \"J7\": 14, \"J8\": 14, \"J9\": 18}</answer>"}
{"model": "anthropic/claude-opus-5.5", "task": "X3", "tier": "extreme", "rep": 0, "pass": true, "detail": "66/66 hidden tests", "routed": "anthropic/claude-opus-5.5", "provider": "Claude Platform on AWS", "cost": 0.030812, "in_tok": 398, "out_tok": 1461, "latency": 14.92, "id": "gen-1791298429-hVO6a3IYgJZRyHpFT6jm", "text": "The search runs Dijkstra over states of the form (row, col, collected-keys mask, opened-doors mask).\n\n- **Why the masks are enough:** there are at most 4 key cells and 4 door cells. The number of keys currently held for a letter is the key cells of that letter already collected minus the doors of that letter already opened. So the two masks capture everything that affects future moves.\n- **State space:** it has at most 100 \u00d7 16 \u00d7 16 = 25,600 states.\n- **Move cost:** entering a digit cell `d` costs `1 + d`, and every other cell costs 1.\n- **Doors:** entering an unopened door requires at least one held key of the matching letter. Opening it consumes that key by setting the door's bit.\n- **Ending at E:** a walk ends on first reaching `E`, so the first time `E` is popped from the queue gives the minimum cost.\n\n```python\nimport heapq\n\ndef vault2(grid: list[str]) -> int:\n    if not grid:\n        return -1\n    R, C = len(grid), len(grid[0])\n    key_id = {}\n    door_id = {}\n    key_letter = []\n    door_letter = []\n    start = None\n    for r in range(R):\n        for c in range(C):\n            ch = grid[r][c]\n            if ch == 'S':\n                start = (r, c)\n            elif ch in 'abc':\n                key_id[(r, c)] = len(key_letter)\n                key_letter.append(ch)\n            elif ch in 'ABC':\n                door_id[(r, c)] = len(door_letter)\n                door_letter.append(ch.lower())\n    if start is None:\n        return -1\n\n    def held(letter, km, dm):\n        cnt = 0\n        for i, l in enumerate(key_letter):\n            if l == letter and (km >> i) & 1:\n                cnt += 1\n        for i, l in enumerate(door_letter):\n            if l == letter and (dm >> i) & 1:\n                cnt -= 1\n        return cnt\n\n    s = (start[0], start[1], 0, 0)\n    dist = {s: 0}\n    pq = [(0, start[0], start[1], 0, 0)]\n    while pq:\n        d, r, c, km, dm = heapq.heappop(pq)\n        if dist.get((r, c, km, dm), float('inf')) < d:\n            continue\n        if grid[r][c] == 'E':\n            return d\n        for dr, dc in ((1, 0), (-1, 0), (0, 1), (0, -1)):\n            nr, nc = r + dr, c + dc\n            if not (0 <= nr < R and 0 <= nc < C):\n                continue\n            ch = grid[nr][nc]\n            if ch == '#':\n                continue\n            cost = 1\n            nkm, ndm = km, dm\n            if ch in '123456789':\n                cost += int(ch)\n            elif ch in 'abc':\n                nkm = km | (1 << key_id[(nr, nc)])\n            elif ch in 'ABC':\n                i = door_id[(nr, nc)]\n                if not (dm >> i) & 1:\n                    if held(ch.lower(), km, dm) <= 0:\n                        continue\n                    ndm = dm | (1 << i)\n            nd = d + cost\n            st = (nr, nc, nkm, ndm)\n            if nd < dist.get(st, float('inf')):\n                dist[st] = nd\n                heapq.heappush(pq, (nd, nr, nc, nkm, ndm))\n    return -1\n```"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "X3", "tier": "extreme", "rep": 0, "pass": true, "detail": "66/66 hidden tests", "routed": "deepseek/deepseek-v4.1-flash", "provider": "AtlasCloud", "cost": 0.005842758, "in_tok": 306, "out_tok": 10283, "latency": 54.13, "id": "gen-1791298390-nVejvarhtKavuE1ELMQd", "text": "```python\ndef vault2(grid: list[str]) -> int:\n    import heapq\n\n    if not grid or not grid[0]:\n        return -1\n\n    R, C = len(grid), len(grid[0])\n    sr = sc = er = ec = -1\n\n    key_cells = []\n    door_cells = []\n    key_idx = [[-1] * C for _ in range(R)]\n    door_idx = [[-1] * C for _ in range(R)]\n\n    for r in range(R):\n        for c, ch in enumerate(grid[r]):\n            if ch == 'S':\n                sr, sc = r, c\n            elif ch == 'E':\n                er, ec = r, c\n            elif ch.islower():\n                key_idx[r][c] = len(key_cells)\n                key_cells.append((r, c, ch))\n            elif ch.isupper() and ch not in ('S', 'E'):\n                door_idx[r][c] = len(door_cells)\n                door_cells.append((r, c, ch))\n\n    if sr == -1 or er == -1:\n        return -1\n\n    if sr == er and sc == ec:\n        return 0\n\n    INF = 10**18\n    dist = {}\n    start = (sr, sc, 0, 0)\n    dist[start] = 0\n    pq = [(0, sr, sc, 0, 0)]\n\n    dirs = [(1, 0), (-1, 0), (0, 1), (0, -1)]\n\n    while pq:\n        cost, r, c, km, dm = heapq.heappop(pq)\n        state = (r, c, km, dm)\n\n        if cost != dist.get(state):\n            continue\n\n        if r == er and c == ec:\n            return cost\n\n        for dr, dc in dirs:\n            nr, nc = r + dr, c + dc\n\n            if nr < 0 or nr >= R or nc < 0 or nc >= C:\n                continue\n\n            ch = grid[nr][nc]\n            if ch == '#':\n                continue\n\n            add = 1\n            if '1' <= ch <= '9':\n                add += int(ch)\n\n            nkm = km\n            ndm = dm\n\n            if key_idx[nr][nc] != -1:\n                idx = key_idx[nr][nc]\n                nkm = km | (1 << idx)\n\n            elif door_idx[nr][nc] != -1:\n                idx = door_idx[nr][nc]\n\n                if ((dm >> idx) & 1) == 0:\n                    letter = door_cells[idx][2].lower()\n\n                    available = 0\n                    for i, (_, _, kch) in enumerate(key_cells):\n                        if kch == letter and ((km >> i) & 1):\n                            available += 1\n\n                    for i, (_, _, dch) in enumerate(door_cells):\n                        if dch.lower() == letter and ((dm >> i) & 1):\n                            available -= 1\n\n                    if available <= 0:\n                        continue\n\n                    ndm = dm | (1 << idx)\n\n            ncost = cost + add\n            nstate = (nr, nc, nkm, ndm)\n\n            if ncost < dist.get(nstate, INF):\n                dist[nstate] = ncost\n                heapq.heappush(pq, (ncost, nr, nc, nkm, ndm))\n\n    return -1\n```"}
{"model": "google/gemini-3.8-flash", "task": "X1", "tier": "extreme", "rep": 0, "pass": true, "detail": "1, 4, 4, -132, 11, 7", "routed": "google/gemini-3.8-flash", "provider": "Google", "cost": 0.036579, "in_tok": 372, "out_tok": 9680, "latency": 54.72, "id": "gen-1791298400-z0u9UaLDG54XtWonzkzP", "text": "<answer>1, 4, 4, -132, 11, 7</answer>"}
{"model": "deepseek/deepseek-v4.1-flash", "task": "X1", "tier": "extreme", "rep": 0, "pass": false, "detail": "", "routed": "deepseek/deepseek-v4.1-flash", "provider": "InferenceNet", "cost": 0.00962737, "in_tok": 391, "out_tok": 32000, "latency": 144.6, "id": "gen-1791298384-oe8qv34gML9tkUMUBoLE", "text": ""}
