Benjamin Schroth
|
d6dcf3b749
|
test(ai-builder): Support evals for agent building and for config evals (no-changelog) (#33888)
Co-authored-by: cubic-dev-ai[bot] <191113872+cubic-dev-ai[bot]@users.noreply.github.com>
|
2026-07-15 12:34:21 +00:00 |
|
José Braulio González Valido
|
2233d720ae
|
ci: Bound and instrument Instance AI eval lane containers (no-changelog) (#33903)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-14 09:28:46 +00:00 |
|
José Braulio González Valido
|
a88231544a
|
feat(ai-builder): Persist eval expectation verdicts to LangSmith run outputs (no-changelog) (#33788)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-09 22:21:46 +00:00 |
|
Anne Aguirre
|
5c6544af4a
|
feat(core): Port agent builder skills and tools to Instance AI (no-changelog) (#33384)
Co-authored-by: Robin Braumann <robin.braumann@n8n.io>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: bjorger <50590409+bjorger@users.noreply.github.com>
Co-authored-by: Michael Drury <michael.drury@n8n.io>
Co-authored-by: Michael Drury <me@michaeldrury.co.uk>
|
2026-07-08 12:41:12 +00:00 |
|
Robin Braumann
|
f25699fe2c
|
test(core): Extend Instance AI intent eval fixtures (no-changelog) (#33224)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-07-07 12:49:32 +00:00 |
|
José Braulio González Valido
|
a89f85b052
|
fix(ai-builder): Make eval verifier resilient to stalls and exclude no-verdict runs from scoring (no-changelog) (#33562)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-07 08:29:53 +00:00 |
|
Milorad FIlipović
|
c5d65998da
|
test: Reduce non-builder noise in MCP workflow evals (no-changelog) (#33565)
|
2026-07-06 13:48:11 +00:00 |
|
José Braulio González Valido
|
a8f1b2c1b8
|
ci: Scope Instance AI evals dispatcher to PR gating (no-changelog) (#33499)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-03 07:29:09 +00:00 |
|
oleg
|
6f0d30fbb1
|
fix: Improve AI Assistant post-build follow-ups (#33419)
|
2026-07-02 11:34:21 +00:00 |
|
José Braulio González Valido
|
b8a993a2ed
|
test(ai-builder): Render full judge text for eval failures, warn only on barely-passed units (no-changelog) (#33127)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-30 09:30:35 +00:00 |
|
José Braulio González Valido
|
c3741ce12a
|
feat(ai-builder): Re-run Instance AI evals against the PR head on demand (no-changelog) (#33148)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-29 06:00:43 +00:00 |
|
José Braulio González Valido
|
d78aba164e
|
test(ai-builder): Add absolute green-gate verdict for pr-tier evals (no-changelog) (#32984)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-26 10:00:06 +00:00 |
|
Milorad FIlipović
|
e12ce0ead0
|
feat(core): Enable LangSmith for mcp evaluations (no-changelog) (#32995)
|
2026-06-25 09:33:22 +00:00 |
|
José Braulio González Valido
|
fe11536df1
|
feat(ai-builder): Add conversation pre-seeding to workflow evals (no-changelog) (#32196)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-06-23 17:30:30 +00:00 |
|
José Braulio González Valido
|
5fa988104b
|
test(ai-builder): Richer multi-turn eval framework (no-changelog) (#32054)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-11 12:57:08 +00:00 |
|
José Braulio González Valido
|
46a222a15b
|
ci: Re-enable Instance AI evals on PRs with cost-bounded trigger (no-changelog) (#31430)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-06-05 08:40:23 +00:00 |
|
José Braulio González Valido
|
700b32237f
|
feat(ai-builder): Surface WHAT-dimension binary checks per built workflow (no-changelog) (#30932)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-26 12:18:52 +01:00 |
|
José Braulio González Valido
|
81ea56fa6b
|
test(ai-builder): Add multi-turn capability for IAI evals (no-changelog) (#30586)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-21 13:03:35 +00:00 |
|
José Braulio González Valido
|
2164afc5df
|
chore(ai-builder): Improve eval comparison alert clarity (no-changelog) (#29929)
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.x) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.14.1) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (25.x) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-06 21:20:49 +00:00 |
|
José Braulio González Valido
|
bbe3e2d148
|
feat(ai-builder): Add per-PR eval regression detection vs LangSmith baseline (#29456)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-06 08:15:08 +00:00 |
|