Commit Graph

100 Commits

Author SHA1 Message Date
José Braulio González Valido
b93a99945b
feat(ai-builder): Harden the eval harness — crash-recovery journal, parity fixes, extension points (no-changelog) (#34747)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 12:25:55 +00:00
José Braulio González Valido
5fd800de20
feat(ai-builder): Support execution scenarios for first-class agents in evals (#34384)
Some checks failed
Build: Benchmark Image / build (push) Waiting to run
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.18.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Util: Sync API Docs / sync-public-api (push) Waiting to run
Test: E2E Performance / prepare-docker (push) Has been cancelled
Release: Storybook / Cloudflare Pages (push) Has been cancelled
Test: E2E Performance / build-and-test-performance (push) Has been cancelled
Test: E2E Performance / Canvas Perf Sentinels (push) Has been cancelled
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 13:54:06 +00:00
José Braulio González Valido
3f7258b1a4
test(ai-builder): Prune the stored-key eval case carve-out (no-changelog) (#34248)
Some checks are pending
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 21:24:40 +00:00
Robin Braumann
cf5ea8e173
fix(core): Make the AI Assistant agent-aware — intent gate, multi-agent sessions, builder session cleanup (no-changelog) (#34200)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-15 13:00:10 +00:00
Arvin A
38894a067a
feat(core): Add config-evals skill for Instance AI eval creation (no-changelog) (#34082)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 12:49:31 +00:00
José Braulio González Valido
ced861603b
test(ai-builder): Move the workflow-eval case corpus to LangTracer (no-changelog) (#33983)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 08:58:49 +00:00
Mutasem Aldmour
199584aab0
fix(core): Correct OpenAI mock response shapes in eval simulation (no-changelog) (#34140)
Co-authored-by: Jose <jose.gonzalez@n8n.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 07:49:16 +00:00
Robin Braumann
fd13039be2
feat(core): Cascade builder sub-agent questions into the AI assistant chat (no-changelog) (#34086)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-14 14:22:03 +00:00
Robin Braumann
f1d156b6cd
feat(core): Route AI assistant agent building through a builder sub-agent (no-changelog) (#34056)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-13 14:52:08 +00:00
Michael Kret
b60aafc809
feat(core): Make Instance AI aware of n8n Connect (#33524)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-07-13 11:14:07 +00:00
Robin Braumann
f40e8e6efd
feat(core): Align Instance AI intent recognition and evals with the anchor + embeds_other taxonomy (no-changelog) (#33843)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: José Braulio González Valido <jose.gonzalez@n8n.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 08:16:46 +00:00
José Braulio González Valido
3bc2292c78
test(ai-builder): Treat answer-only cases as valid without a built workflow (no-changelog) (#33942)
Some checks are pending
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 21:54:14 +00:00
José Braulio González Valido
13c95f2c05
fix(ai-builder): Eliminate non-builder noise families in workflow evals (no-changelog) (#33944)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 13:25:29 +00:00
José Braulio González Valido
7dd065a77b
feat(ai-builder): Scale eval budgets by case complexity and bound the heaviest prompts (no-changelog) (#33950)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 12:58:17 +00:00
Arvin A
e6dee743e6
test(ai-builder): Add advanced user-behavior multi-turn eval cases (no-changelog) (#31758)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: José Braulio González Valido <jose.gonzalez@n8n.io>
Co-authored-by: José Braulio González Valido <josebragv@gmail.com>
2026-07-09 18:13:42 +00:00
oleg
a1df115aef
fix(core): Report auto-bound credentials from AI Assistant workflow builds (#33802) 2026-07-08 13:01:29 +00:00
oleg
2d42b3fecf
fix(core): Reconcile stale AI Assistant mocked-credential simulation plan before workflow verification (#33808) 2026-07-08 13:01:08 +00:00
Anne Aguirre
5c6544af4a
feat(core): Port agent builder skills and tools to Instance AI (no-changelog) (#33384)
Co-authored-by: Robin Braumann <robin.braumann@n8n.io>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: bjorger <50590409+bjorger@users.noreply.github.com>
Co-authored-by: Michael Drury <michael.drury@n8n.io>
Co-authored-by: Michael Drury <me@michaeldrury.co.uk>
2026-07-08 12:41:12 +00:00
Benjamin Schroth
48e7900db0
test(ai-builder): Add eval cases for evaluation generation (#33719)
Some checks are pending
Build: Benchmark Image / build (push) Waiting to run
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Util: Sync API Docs / sync-public-api (push) Waiting to run
2026-07-07 15:29:42 +00:00
Albert Alises
e83512dfd8
test(core): Add Instance AI eval cases for recently merged builder changes (no-changelog) (#33763) 2026-07-07 14:12:08 +00:00
Robin Braumann
f25699fe2c
test(core): Extend Instance AI intent eval fixtures (no-changelog) (#33224)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 12:49:32 +00:00
Riqwan Thamir
0775ea114a
refactor(core): Remove Instance AI orchestration delegate tool (no-changelog) (#33556)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 09:15:26 +00:00
Albert Alises
3aa3ffd13b
fix(core): Report node errors in workflow verification (#33576) 2026-07-07 08:21:38 +00:00
Riqwan Thamir
d21a499462
test(core): Add Instance AI loop pattern workflow eval cases (no-changelog) (#33712)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 07:26:51 +00:00
Milorad FIlipović
c5d65998da
test: Reduce non-builder noise in MCP workflow evals (no-changelog) (#33565) 2026-07-06 13:48:11 +00:00
José Braulio González Valido
4d11639cf1
test(ai-builder): Add plan-rejection eval and harden multi-turn eval consistency (no-changelog) (#33433)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 06:11:22 +00:00
Albert Alises
2be4092c4d
fix(core): Improve Instance AI workflow-building reliability (#33530) 2026-07-03 13:27:45 +00:00
oleg
6f0d30fbb1
fix: Improve AI Assistant post-build follow-ups (#33419) 2026-07-02 11:34:21 +00:00
Albert Alises
4192c9554e
feat(core): Prefer predefined credentials over generic auth on HTTP Request nodes (#33298) 2026-07-02 07:50:49 +00:00
José Braulio González Valido
2ad33a1303
fix(ai-builder): Fix AI-node eval mock-execution timeouts (no-changelog) (#33323)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-01 08:33:51 +00:00
José Braulio González Valido
65f375bbd4
feat(ai-builder): Make eval executionScenarios optional for build-only cases (no-changelog) (#33252)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 09:41:14 +00:00
José Braulio González Valido
7ea6900e58
feat(ai-builder): Honor seedThread.liveTurnRunId in eval reconstruction (no-changelog) (#33251)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 09:34:10 +00:00
Mutasem Aldmour
b67ef426e8
test: Add Instance AI flight-status-change workflow eval (#33228)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 06:55:40 +00:00
José Braulio González Valido
1d28330d18
feat(ai-builder): Dual-tenant LangSmith reads for Instance AI evals (no-changelog) (#33230)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 17:17:14 +00:00
Riqwan Thamir
6d550faf7d
chore(core): Add test on voice based agent (#33226) 2026-06-29 14:23:40 +00:00
José Braulio González Valido
23d70d1bb9
feat(ai-builder): Source eval test cases from LangTracer (no-changelog) (#33067)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 10:17:49 +00:00
Albert Alises
de8f2984ed
fix(core): Stop the builder from re-running a workflow it just verified (no-changelog) (#33141) 2026-06-29 08:10:44 +00:00
Jaakko Husso
691d39d9c8
fix(core): Ground WhatsApp Trigger verify-token guidance and add regression evals (no-changelog) (#33095)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 07:39:08 +00:00
José Braulio González Valido
1f012f8cb5
test(ai-builder): Reproduce explicit scenario values in eval mocks (no-changelog) (#33166)
Some checks failed
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Util: Update Node Popularity / update-popularity (push) Has been cancelled
Util: Update Node Popularity / approve-and-automerge (push) Has been cancelled
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 15:15:15 +00:00
oleg
a00ac2722c
fix(core): Make AI Assistant workflow verification and setup repeatable (#33085)
Signed-off-by: Oleg Ivaniv <me@olegivaniv.com>
2026-06-26 13:39:28 +00:00
Jaakko Husso
3b27616b41
fix(core): Settle failed workflow verification instead of looping (no-changelog) (#33111)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 13:14:04 +00:00
José Braulio González Valido
d78aba164e
test(ai-builder): Add absolute green-gate verdict for pr-tier evals (no-changelog) (#32984)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 10:00:06 +00:00
Albert Alises
9161865a72
fix(core): Improve builder handling of errored expressions (no-changelog) (#33059) 2026-06-26 09:23:39 +00:00
Jaakko Husso
713203a1c5
fix(core): Surface infeasible core capabilities instead of silently downgrading (no-changelog) (#33063)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 09:19:17 +00:00
Albert Alises
cf8ef2b1b7
test(core): Add eval ensuring builder decides technical choices instead of asking (no-changelog) (#33053)
Some checks failed
Build: Benchmark Image / build (push) Waiting to run
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Util: Sync API Docs / sync-public-api (push) Waiting to run
Release: Schedule Patch Release PRs / Create patch release PR (${{ matrix.track }}) (beta) (push) Has been cancelled
Release: Schedule Patch Release PRs / Create patch release PR (${{ matrix.track }}) (stable) (push) Has been cancelled
Release: Schedule Patch Release PRs / Create patch release PR (${{ matrix.track }}) (v1) (push) Has been cancelled
2026-06-25 22:33:38 +00:00
Albert Alises
02df832dd2
fix: Prefer placeholders over pre-build setup questions (#32969)
Some checks are pending
Build: Benchmark Image / build (push) Waiting to run
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Util: Sync API Docs / sync-public-api (push) Waiting to run
2026-06-25 17:39:39 +00:00
Riqwan Thamir
58fb6ca742
feat(core): Add more use case evaluations to instance ai (no-changelog) (#32963) 2026-06-25 14:15:30 +00:00
Milorad FIlipović
5b77a717b0
fix(core): Handle array-style conversations in mcp evals (no-changelog) (#33013) 2026-06-25 11:53:57 +00:00
Riqwan Thamir
b3a399553f
fix(core): Broken workflow verification loop in iAI (no-changelog) (#32979)
Co-authored-by: Oleg Ivaniv <me@olegivaniv.com>
2026-06-25 08:20:46 +00:00
Milorad FIlipović
d3dd105aae
feat(core): Make instanceAI evals more mcp-friendly (no-changelog) (#32916) 2026-06-25 06:33:50 +00:00