Last 12 weeks · 0 commits
2 of 6 standards met
@dummdidumm @paoloricciuti @teemingc @khromov @ghostdevv We discussed a kind of matrix of capability assessments ranging from 'do you know this API' through to 'can you decompose this task and translate it to svelte'. We also discussed environments. Here are some initial ideas. Capabilities Task style Best use API-directed API/syntax capability Behaviour-directed Primary capability eval Constraint-directed Useful middle ground Debugging Realistic capability eval API debugging Testing mastery of a particular API Compositional Multi-capability eval Vertical Agent/system capability Hanresses light weight (plain Pi) with Svelte MCP big boy Environments blank simple template existing repo Svelte/ kit coverage Initial list: Svelte | SvelteKit | -- State | Routing/layouts Derived state | Universal/server load Effects | Data dependency/invalidation Props | Server/client boundaries Bindings | SSR + hydration Events | Navigation Lists/identity | Form actions Conditional/async rendering | Progressive enhancement Snippets/composition | Remote functions Context | Endpoints/HTTP Lifecycle | Authentication/cookies Reusable reactive logic | App/server state Forms | Errors/redirects TypeScript | Environment/secrets SSR safety | Hooks Accessibility | Rendering modes Debugging/compiler diagnostics | Generated types Legacy Svelte interoperability | Deployment boundaries Which gives us a matrix that looks something like this for each capability: Which might be something like: That is 27 evals per 'capability', what is nice about this is we already have most of the core 'api' directed capabilities in the repo, the behavioural flavour is mostly a different prompt. debugging could also be a slight modification. For more 'vertical' challenges ('create a payment cart that updates live') we don't need one per api slice, because they naturally cover multiple APIs at once. The real thing we are evaluating there is 'can this problem be decomposed and traced back to the correct Svelte APIs?'. Which brings us to evaluations We discussed correctness vs quality. There are probably other axes too but the nice thing about this is we don't need more task samples, instead we need additional assessment criteria. Every example could be checked for 'does it work', 'does the code use appropriate patterns', 'is the code accessible', 'does the code fall into any performance traps'. Some of these are deterministically verifiable, some of them aren't but we can discussed verification more elsewhere.
[x] Review repository scripts and existing eval/experiment structure [x] Draft README instructions for running experiments and visualizing results [x] Document how to create new evals and experiments [x] Verify documentation accuracy against current scripts/structure 💬 We'd love your input! Share your thoughts on Copilot coding agent in our 2 minute survey.
Repository: sveltejs/svelte-evals. Description: Evals for LLMS to learn/benchmark their Svelte skills Stars: 16, Forks: 2. Primary language: TypeScript. Languages: TypeScript (77.5%), JavaScript (18.8%), HTML (2.8%), Svelte (1%). License: MIT. Open PRs: 0, open issues: 3. Last activity: 6mo ago. Community health: 62%. Top contributors: paoloricciuti, Copilot, dummdidumm.