GitShow/sveltejs/svelte-evals
sveltejs

svelte-evals

Evals for LLMS to learn/benchmark their Svelte skills

by sveltejs
Star on GitHubForknpm

TypeScript

16 stars2 forks3 contributorsQuiet · 6mo agoSince 2025MIT

Meet the team

See all 3 on GitHub →
paoloricciuti
paoloricciuti11 contributions
CopilotBot
Copilot3 contributions
dummdidumm
dummdidumm2 contributions

Languages

View on GitHub →
TypeScript77.5%
JavaScript18.8%
HTML2.8%
Svelte1%

Commit activity

Last 12 weeks · 0 commits

Full graph →

Community health

2 of 6 standards met

Community profile →
62
✓README✓License○Contributing○Code of Conduct○Issue Template○PR Template

Recent PRs & issues

Quiet · 3 discussions · Last activity 6mo ago
See all on GitHub →
pngwn
Eval CapabilitiesOpenIssue

@dummdidumm @paoloricciuti @teemingc @khromov @ghostdevv We discussed a kind of matrix of capability assessments ranging from 'do you know this API' through to 'can you decompose this task and translate it to svelte'. We also discussed environments. Here are some initial ideas. Capabilities Task style Best use API-directed API/syntax capability Behaviour-directed Primary capability eval Constraint-directed Useful middle ground Debugging Realistic capability eval API debugging Testing mastery of a particular API Compositional Multi-capability eval Vertical Agent/system capability Hanresses light weight (plain Pi) with Svelte MCP big boy Environments blank simple template existing repo Svelte/ kit coverage Initial list: Svelte | SvelteKit | -- State | Routing/layouts Derived state | Universal/server load Effects | Data dependency/invalidation Props | Server/client boundaries Bindings | SSR + hydration Events | Navigation Lists/identity | Form actions Conditional/async rendering | Progressive enhancement Snippets/composition | Remote functions Context | Endpoints/HTTP Lifecycle | Authentication/cookies Reusable reactive logic | App/server state Forms | Errors/redirects TypeScript | Environment/secrets SSR safety | Hooks Accessibility | Rendering modes Debugging/compiler diagnostics | Generated types Legacy Svelte interoperability | Deployment boundaries Which gives us a matrix that looks something like this for each capability: Which might be something like: That is 27 evals per 'capability', what is nice about this is we already have most of the core 'api' directed capabilities in the repo, the behavioural flavour is mostly a different prompt. debugging could also be a slight modification. For more 'vertical' challenges ('create a payment cart that updates live') we don't need one per api slice, because they naturally cover multiple APIs at once. The real thing we are evaluating there is 'can this problem be decomposed and traced back to the correct Svelte APIs?'. Which brings us to evaluations We discussed correctness vs quality. There are probably other axes too but the nice thing about this is we don't need more task samples, instead we need additional assessment criteria. Every example could be checked for 'does it work', 'does the code use appropriate patterns', 'is the code accessible', 'does the code fall into any performance traps'. Some of these are deterministically verifiable, some of them aren't but we can discussed verification more elsewhere.

pngwn · 1w ago
karimfromjordan
Case for 'using a vanilla JS lib'OpenIssue

It might be worth adding multiple cases because it depends on what exactly the lib does. but the most obvious example would be a lib like Chart.js which renders charts into a specific HTML element. In Svelte 5 this should probably be done with an Attachment.

karimfromjordan · 1mo ago
teemingc
Case for "passing state upwards"OpenIssue

This one's kind of open-ended. I'd like to see the Svelte-way for a child component passing state to a parent component. The answer is usually "don't", but when LLMs are React-trained and we don't one-to-one APIs, such as a portal-ing something, I'd like to see what the recommended approach is.

teemingc · 1mo ago

Recent fixes

View closed PRs →
paoloricciuti
feat: bring old svelte bench test overMergedPR
paoloricciuti · 7mo ago
Copilot
Document agent-eval experiment workflow and scaffoldingMergedPR

[x] Review repository scripts and existing eval/experiment structure [x] Draft README instructions for running experiments and visualizing results [x] Document how to create new evals and experiments [x] Verify documentation accuracy against current scripts/structure 💬 We'd love your input! Share your thoughts on Copilot coding agent in our 2 minute survey.

Copilot · 7mo ago
Structured data for AI agents

Repository: sveltejs/svelte-evals. Description: Evals for LLMS to learn/benchmark their Svelte skills Stars: 16, Forks: 2. Primary language: TypeScript. Languages: TypeScript (77.5%), JavaScript (18.8%), HTML (2.8%), Svelte (1%). License: MIT. Open PRs: 0, open issues: 3. Last activity: 6mo ago. Community health: 62%. Top contributors: paoloricciuti, Copilot, dummdidumm.

·@ofershap

Replace github.com with gitshow.dev