José Braulio González Valido
b93a99945b
feat(ai-builder): Harden the eval harness — crash-recovery journal, parity fixes, extension points (no-changelog) ( #34747 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 12:25:55 +00:00
José Braulio González Valido
5fd800de20
feat(ai-builder): Support execution scenarios for first-class agents in evals ( #34384 )
...
Build: Benchmark Image / build (push) Waiting to run
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.18.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Util: Sync API Docs / sync-public-api (push) Waiting to run
Test: E2E Performance / prepare-docker (push) Has been cancelled
Release: Storybook / Cloudflare Pages (push) Has been cancelled
Test: E2E Performance / build-and-test-performance (push) Has been cancelled
Test: E2E Performance / Canvas Perf Sentinels (push) Has been cancelled
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 13:54:06 +00:00
José Braulio González Valido
3f7258b1a4
test(ai-builder): Prune the stored-key eval case carve-out (no-changelog) ( #34248 )
...
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 21:24:40 +00:00
Robin Braumann
cf5ea8e173
fix(core): Make the AI Assistant agent-aware — intent gate, multi-agent sessions, builder session cleanup (no-changelog) ( #34200 )
...
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-15 13:00:10 +00:00
Arvin A
38894a067a
feat(core): Add config-evals skill for Instance AI eval creation (no-changelog) ( #34082 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 12:49:31 +00:00
José Braulio González Valido
ced861603b
test(ai-builder): Move the workflow-eval case corpus to LangTracer (no-changelog) ( #33983 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 08:58:49 +00:00
Mutasem Aldmour
199584aab0
fix(core): Correct OpenAI mock response shapes in eval simulation (no-changelog) ( #34140 )
...
Co-authored-by: Jose <jose.gonzalez@n8n.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 07:49:16 +00:00
Robin Braumann
fd13039be2
feat(core): Cascade builder sub-agent questions into the AI assistant chat (no-changelog) ( #34086 )
...
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-14 14:22:03 +00:00
Robin Braumann
f1d156b6cd
feat(core): Route AI assistant agent building through a builder sub-agent (no-changelog) ( #34056 )
...
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-13 14:52:08 +00:00
Michael Kret
b60aafc809
feat(core): Make Instance AI aware of n8n Connect ( #33524 )
...
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-07-13 11:14:07 +00:00
Robin Braumann
f40e8e6efd
feat(core): Align Instance AI intent recognition and evals with the anchor + embeds_other taxonomy (no-changelog) ( #33843 )
...
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: José Braulio González Valido <jose.gonzalez@n8n.io>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 08:16:46 +00:00
José Braulio González Valido
3bc2292c78
test(ai-builder): Treat answer-only cases as valid without a built workflow (no-changelog) ( #33942 )
...
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 21:54:14 +00:00
José Braulio González Valido
13c95f2c05
fix(ai-builder): Eliminate non-builder noise families in workflow evals (no-changelog) ( #33944 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 13:25:29 +00:00
José Braulio González Valido
7dd065a77b
feat(ai-builder): Scale eval budgets by case complexity and bound the heaviest prompts (no-changelog) ( #33950 )
...
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 12:58:17 +00:00
Arvin A
e6dee743e6
test(ai-builder): Add advanced user-behavior multi-turn eval cases (no-changelog) ( #31758 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: José Braulio González Valido <jose.gonzalez@n8n.io>
Co-authored-by: José Braulio González Valido <josebragv@gmail.com>
2026-07-09 18:13:42 +00:00
oleg
a1df115aef
fix(core): Report auto-bound credentials from AI Assistant workflow builds ( #33802 )
2026-07-08 13:01:29 +00:00
oleg
2d42b3fecf
fix(core): Reconcile stale AI Assistant mocked-credential simulation plan before workflow verification ( #33808 )
2026-07-08 13:01:08 +00:00
Anne Aguirre
5c6544af4a
feat(core): Port agent builder skills and tools to Instance AI (no-changelog) ( #33384 )
...
Co-authored-by: Robin Braumann <robin.braumann@n8n.io>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: bjorger <50590409+bjorger@users.noreply.github.com>
Co-authored-by: Michael Drury <michael.drury@n8n.io>
Co-authored-by: Michael Drury <me@michaeldrury.co.uk>
2026-07-08 12:41:12 +00:00
Benjamin Schroth
48e7900db0
test(ai-builder): Add eval cases for evaluation generation ( #33719 )
Build: Benchmark Image / build (push) Waiting to run
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Util: Sync API Docs / sync-public-api (push) Waiting to run
2026-07-07 15:29:42 +00:00
Albert Alises
e83512dfd8
test(core): Add Instance AI eval cases for recently merged builder changes (no-changelog) ( #33763 )
2026-07-07 14:12:08 +00:00
Robin Braumann
f25699fe2c
test(core): Extend Instance AI intent eval fixtures (no-changelog) ( #33224 )
...
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 12:49:32 +00:00
Riqwan Thamir
0775ea114a
refactor(core): Remove Instance AI orchestration delegate tool (no-changelog) ( #33556 )
...
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 09:15:26 +00:00
Albert Alises
3aa3ffd13b
fix(core): Report node errors in workflow verification ( #33576 )
2026-07-07 08:21:38 +00:00
Riqwan Thamir
d21a499462
test(core): Add Instance AI loop pattern workflow eval cases (no-changelog) ( #33712 )
...
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 07:26:51 +00:00
Milorad FIlipović
c5d65998da
test: Reduce non-builder noise in MCP workflow evals (no-changelog) ( #33565 )
2026-07-06 13:48:11 +00:00
José Braulio González Valido
4d11639cf1
test(ai-builder): Add plan-rejection eval and harden multi-turn eval consistency (no-changelog) ( #33433 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 06:11:22 +00:00
Albert Alises
2be4092c4d
fix(core): Improve Instance AI workflow-building reliability ( #33530 )
2026-07-03 13:27:45 +00:00
oleg
6f0d30fbb1
fix: Improve AI Assistant post-build follow-ups ( #33419 )
2026-07-02 11:34:21 +00:00
Albert Alises
4192c9554e
feat(core): Prefer predefined credentials over generic auth on HTTP Request nodes ( #33298 )
2026-07-02 07:50:49 +00:00
José Braulio González Valido
2ad33a1303
fix(ai-builder): Fix AI-node eval mock-execution timeouts (no-changelog) ( #33323 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-01 08:33:51 +00:00
José Braulio González Valido
65f375bbd4
feat(ai-builder): Make eval executionScenarios optional for build-only cases (no-changelog) ( #33252 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 09:41:14 +00:00
José Braulio González Valido
7ea6900e58
feat(ai-builder): Honor seedThread.liveTurnRunId in eval reconstruction (no-changelog) ( #33251 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 09:34:10 +00:00
Mutasem Aldmour
b67ef426e8
test: Add Instance AI flight-status-change workflow eval ( #33228 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 06:55:40 +00:00
José Braulio González Valido
1d28330d18
feat(ai-builder): Dual-tenant LangSmith reads for Instance AI evals (no-changelog) ( #33230 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 17:17:14 +00:00
Riqwan Thamir
6d550faf7d
chore(core): Add test on voice based agent ( #33226 )
2026-06-29 14:23:40 +00:00
José Braulio González Valido
23d70d1bb9
feat(ai-builder): Source eval test cases from LangTracer (no-changelog) ( #33067 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 10:17:49 +00:00
Albert Alises
de8f2984ed
fix(core): Stop the builder from re-running a workflow it just verified (no-changelog) ( #33141 )
2026-06-29 08:10:44 +00:00
Jaakko Husso
691d39d9c8
fix(core): Ground WhatsApp Trigger verify-token guidance and add regression evals (no-changelog) ( #33095 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 07:39:08 +00:00
José Braulio González Valido
1f012f8cb5
test(ai-builder): Reproduce explicit scenario values in eval mocks (no-changelog) ( #33166 )
...
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Util: Update Node Popularity / update-popularity (push) Has been cancelled
Util: Update Node Popularity / approve-and-automerge (push) Has been cancelled
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 15:15:15 +00:00
oleg
a00ac2722c
fix(core): Make AI Assistant workflow verification and setup repeatable ( #33085 )
...
Signed-off-by: Oleg Ivaniv <me@olegivaniv.com>
2026-06-26 13:39:28 +00:00
Jaakko Husso
3b27616b41
fix(core): Settle failed workflow verification instead of looping (no-changelog) ( #33111 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 13:14:04 +00:00
José Braulio González Valido
d78aba164e
test(ai-builder): Add absolute green-gate verdict for pr-tier evals (no-changelog) ( #32984 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 10:00:06 +00:00
Albert Alises
9161865a72
fix(core): Improve builder handling of errored expressions (no-changelog) ( #33059 )
2026-06-26 09:23:39 +00:00
Jaakko Husso
713203a1c5
fix(core): Surface infeasible core capabilities instead of silently downgrading (no-changelog) ( #33063 )
...
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 09:19:17 +00:00
Albert Alises
cf8ef2b1b7
test(core): Add eval ensuring builder decides technical choices instead of asking (no-changelog) ( #33053 )
Build: Benchmark Image / build (push) Waiting to run
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Util: Sync API Docs / sync-public-api (push) Waiting to run
Release: Schedule Patch Release PRs / Create patch release PR (${{ matrix.track }}) (beta) (push) Has been cancelled
Release: Schedule Patch Release PRs / Create patch release PR (${{ matrix.track }}) (stable) (push) Has been cancelled
Release: Schedule Patch Release PRs / Create patch release PR (${{ matrix.track }}) (v1) (push) Has been cancelled
2026-06-25 22:33:38 +00:00
Albert Alises
02df832dd2
fix: Prefer placeholders over pre-build setup questions ( #32969 )
Build: Benchmark Image / build (push) Waiting to run
CI: Master (Build, Test, Lint) / Build for Github Cache (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (22.22.3) (push) Waiting to run
CI: Master (Build, Test, Lint) / Unit tests (24.16.0) (push) Waiting to run
CI: Master (Build, Test, Lint) / Lint (push) Waiting to run
CI: Master (Build, Test, Lint) / Performance (push) Waiting to run
CI: Master (Build, Test, Lint) / Notify Slack on failure (push) Blocked by required conditions
Util: Sync API Docs / sync-public-api (push) Waiting to run
2026-06-25 17:39:39 +00:00
Riqwan Thamir
58fb6ca742
feat(core): Add more use case evaluations to instance ai (no-changelog) ( #32963 )
2026-06-25 14:15:30 +00:00
Milorad FIlipović
5b77a717b0
fix(core): Handle array-style conversations in mcp evals (no-changelog) ( #33013 )
2026-06-25 11:53:57 +00:00
Riqwan Thamir
b3a399553f
fix(core): Broken workflow verification loop in iAI (no-changelog) ( #32979 )
...
Co-authored-by: Oleg Ivaniv <me@olegivaniv.com>
2026-06-25 08:20:46 +00:00
Milorad FIlipović
d3dd105aae
feat(core): Make instanceAI evals more mcp-friendly (no-changelog) ( #32916 )
2026-06-25 06:33:50 +00:00