GitShow/jaredpalmer/kev
jaredpalmer

kev

Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own

by jaredpalmer
decision-modeljevqwen3
Star on GitHubFork

Python

8.0k stars504 forks10 contributorsActive · 21h agoSince 2026kev-familyApache-2.0

Meet the team

See all 10 on GitHub →
jaredpalmer
jaredpalmer286 contributions
devin-ai-integration[bot]Bot
devin-ai-integration[bot]12 contributions
lunar-me
lunar-me2 contributions
bhimrazy
bhimrazy1 contribution
TwelveNights
TwelveNights1 contribution
ImgBotApp
ImgBotApp1 contribution
Radexito
Radexito1 contribution
VecSzn
VecSzn1 contribution

Languages

View on GitHub →
Python95.1%
TypeScript4.3%
CSS0.3%
Shell0.2%
JavaScript0.1%
HTML0%

Commit activity

Last 12 weeks · 307 commits

Full graph →

Community health

2 of 6 standards met

Community profile →
42
✓README✓License○Contributing○Code of Conduct○Issue Template○PR Template

Recent PRs & issues

Active · Last activity 21h ago
See all on GitHub →
Prajeeth-12
fix: warn at import time when wrong PyPI 'kev' package is installedOpenPR

Problem PyPI already has an unrelated package named . Anyone who runs silently installs the wrong thing, then gets mysterious s when they try to use this project. Reported in #159. Fix Add a runtime guard in (previously empty). On every it checks the installed distribution's metadata. If the / fields don't contain , it fires an with a clear message pointing to the correct install command: This uses only (stdlib, Python 3.8+), adds no dependencies, and is a no-op when the right package is installed. The makes it safe if metadata is missing for any reason. What was not changed is left as-is — renaming the distribution is a separate decision that requires coordination around a PyPI publish under a new slug (e.g. ). That can be a follow-up if desired. Closes #159

Prajeeth-12 · 16h ago
k-o-n-t-o-r
KEV_QUANT / KEV_BASE: serve Kev-27B from an int8 or nf4 base on one 40 GB GPUOpenPR

Kev-27B needs one H100 80 GB or an H200 today: its frozen Qwen3.8-27B backbone is 55 GB resident in bf16, and nothing smaller can load it. The part Kev trained is small (the LoRA adapter is 0.5 GB, the pointer head 10 MB); the memory is the base. This PR lets the base load quantized, int8 or 4-bit NF4, so Kev-27B runs on a 40 GB A100 (int8, 25.3 GiB loaded) or, by the numbers below, a 24 GB card (nf4, 14.3 GiB). Mechanism () passes a quantization config to the base's , built in one place, : int8 is torchao weight-only: int8 weights with one bf16 scale per output channel, activations stay bf16 and weights are dequantized inside each matmul. nf4 is bitsandbytes 4-bit NF4 (blocks of 64, double-quantized scales, bf16 compute). The adapter stays unmerged: LoRA adds its low-rank update next to the quantized instead of being folded into . That is already how a bf16-backbone checkpoint loads when is off; quantization now forces it (), because folding would requantize . An int8 row's step is , and most elements of a rank-16 delta are smaller than half a step, so they would round away. Since fused kernels need merged weights, is off for quantized loads too. Two kinds of weight stay bf16 (): , which Kev does not use. The Gated DeltaNet decay and beta projections, / . They are tiny (hidden x 48), and their output multiplies the recurrent state at every token, so an error there compounds over the whole state. The skip list is a regex, , because the two libraries match different names: torchao matches parameter names (), bitsandbytes module names (). A bare works for bitsandbytes and silently does nothing for torchao. My first int8 run had exactly that bug and quantized both; the int8 row below is the corrected run. () swaps in a base saved pre-quantized by . sees in its and loads it as is instead of quantizing again. That is a 27.5 GiB (int8) or 16.5 GiB (nf4) download instead of 52 GiB, with nothing to quantize at load. The saved base carries the original's tokenizer (identical encodings, delimiter ids and vocab), and loads both from it, so an offline run does not need the original base cached. A full-weight checkpoint has no base to replace and refuses . requires , because the precision decides the loader's device map and the no-merge rule. checks, on the tiny hybrid base: the saved and the at-load base give bit-identical probabilities; stays a bf16 while is quantized; the adapter is still a even with . Why not bitsandbytes LLM.int8 I tried it first (8-bit bitsandbytes). It keeps outlier columns in fp16, but it picks those columns from the whole batch. So a question's answer depends on which other requests share its batch. That breaks the isolation Kev serves under: a question's answer must not depend on what else is asked with it. On the served path, CUDA-graph batches against eager moved answers by up to 0.040, against bf16's 0.009. CUDA-graph capture also crashed inside it (a assert), and it was 2.6x slower than bf16. torchao weight-only has none of these problems: served graphs vs eager is 0.008. Results Checkpoint and splits: on the transfer-v4 and decision-v7 development partitions (764 and 1,468 questions; accuracy and Brier over the 656 and 1,264 clean ones), scored by 's path (). Reference: bf16 on an RTX PRO 6000 (Colab G4). It reproduces the model card's transfer-v4 dev 0.848 / 0.229. int8 and nf4: the pre-quantized bases loaded back through on an A100 40GB with . : a question's largest probability change against bf16. A flip is a changed argmax. Other GPU noise:** the A100 rows also carry the arithmetic differences between the two GPUs. No decision rule was registered before these runs, so I read them against the nearest bar the repo has: the release tolerance for bf16 serving against fp32 (max =3.7.1modal_app.pypyprojectuv.locktorch.compileskills/kev-deployscripts/save_quantized.pypythontest_quantized_base_keeps_the_adapter_unmergedtests/test_conventions.pyKEV_QUANTKEV_BASELoadOptions.from_envTorchAoConfigBitsAndBytesConfigkev.modelscripts/quant_eval.pyKEV_BASEscripts/quant_compare.pytests/test_model.pyquantbase`).

k-o-n-t-o-r · 17h ago
jaredpalmer
Release CANDIDATE (private, do not release yet): Kev-27B v2 = round 23's 27b-k-w85 staged on the volume at T 1.32, private Hub repo jaredpalmer/kev-27b-v2-candidate, model cardOpenPR

Release candidate, not a release. Do not merge before Jared decides. Nothing in this PR or the steps behind it is public. There is no public Hub repo or revision, no public card, and no README, Space, collection or GitHub-release change. Stacked on #176 (round 23 result and confirmation, branch ), whose committed verdicts this card cites. Round 23 confirmed one checkpoint, : round 22's final full-weight SFT of Qwen3.8-27B, blended 0.85 / 0.15 toward the released Kev-27B, with the SFT's pointer head. It passed every registered stage against Kev-27B: breadth-v1 test +1.2 pp [+0.3, +2.2]; tasksource-heldout-v1 test +5.3 [+3.7, +6.8]; pooled hard / devtools / documents-v1 test +8.9 [+7.5, +10.3]; locked transfer-v4 0.8887 (583 of 656) against the 0.886 bar, with served Brier 0.154; the bf16 serving check. Until now the checkpoint existed only as a round arm on the Modal volume. Its still carried the raw temperature 1.0. The release temperature had been fitted only on a local copy of , and there was no card. This PR turns the arm into a reviewable private candidate, Kev-27B v2. It keeps every step reproducible and leaves the round's own files untouched. What was done, in order (record: and ) 1. Copy on the volume, never over anything. (new, ) runs in a CPU container under app . It copied to , and it refuses a destination that exists. It hashes both copies the way does (shards hashed in parallel, same value). Both are , the round's hash. The source is untouched (, T 1.0). Naming: the internal id is , because already names the released Kev-27B (B1 v2) in , in and on the volume. Only the display name is "Kev-27B v2". 2. Release temperature. ran on the copy's with the round's registered pool rows: , limited to its eight held-out sources; , limited to ; minus the records; 648 questions in all. It found no overlap with training and fitted T = 1.3195, the round's pool fit exactly. Out-of-fold ECE goes from 0.048 to 0.038 and the intervals overlap; the 90 % interval of T is [1.20, 1.45]. The resulting () is byte-identical to the monitor's local fit, and it went back into the copy only. 3. Private Hub repo. (new) runs in a CPU container, so the 51 GB never touch a laptop. It created as a private repo and uploaded commit . The card README was updated later, at ; the weights have not changed since . now refuses an existing public repo (). Before, it passed to , which does nothing to a repo that already exists, so a candidate could have landed in a public repo. It links full-weight shards into its staging directory instead of copying them. It uploads with the trial files. The Hub's LFS hashes give the same weights hash () and hash (). 4. The Hub copy loads and scores exactly. ran on an H200, loading with a token. All 252 rows were served at T 1.3195. Their logits equal round 23's committed raw logits times fp32(1/T) bit for bit (PyTorch divides by a scalar as a multiply by its reciprocal), so the backbone and head outputs are identical. Argmax matches on 252 of 252. 5. Card and numbers. follows the formal card style, and its 219 numbers are in . Sources: the round's verdicts and read-out, the test breadth report, the serving reports, the staging record and . comes from . The script now accepts an arm with a registered in place of a trial, and prints where a pool arm has no held-out-pairs figure. For recorded releases whose rows are missing, the Jev documents figure and the locked decision read are now optional. All four recorded releases (, , , ) reproduce byte for byte from the research checkout's rows. What the card says plainly. These points are in "Read this first", not buried: Post-hoc selection. Round 23 was designed after round 24's confirmation and confirmed on the same test and locked partitions. Short states are not better than Kev-27B. Locked −0.8 [−2.0, +0.5]. On transfer-r3 test with included, −2.1 [−3.5, −0.8]. Long contracts are worse. CUAD in longdoc-v1 test scores −1.6 pp [−3.0, −0.3], with ECE 0.053 against 0.007. That read landed while this was staged and is in the card. No scienthoon read. The suite was removed before round 23; the SFT parent was −5.5 pp there. In distribution. hard-v1, devtools-v1 and documents-v1 are in the training data. The tasksource family list stays private ("119 commercially licensed task families; list private"). , not . Every weight derives from fine-tunes of the one base, and the Hub's is for merges of several listed base models. Kev-27B is itself an adapter of the same base. Not in this PR Releasing. If Jared approves: 1. tag Kev-27B's current weights; 2. publish the staged copy to from a container; 3. move the card to ; 4. then update the README, the Space, the collection and kev-deploy. Spend. Modal metered $3,971.71 before and about $3,977 after (workspace-wide, other apps running). App : a CPU copy and upload plus one ~10-minute H200 read, about $1-2. Test plan [x] Unit suites (the CI job) on this branch: 452 passed, 22 skipped, 1 failed. The failure was a without in the new tests (); after the fix, passes (19) and so do the two new tests, and . [x] : 978 claims verified. [x] : the four recorded releases are byte-identical to their committed JSON, run from the research checkout's rows. [x] Volume copy: weights sha256 equal (). Source untouched. The copy's is at T 1.3195. [x] Hub: the repo is private; LFS-derived weights and hashes are equal; semif-v1 rows from the Hub load match round 23's read, 252 of 252 bit for bit before T.

jaredpalmer · 1d ago

Recent fixes

View closed PRs →
TwelveNights
docs: allow for ROCm 7.14 installation extraMergedPR

As a follow-up to https://github.com/jaredpalmer/kev/pull/23, this allows an installer to automatically fetch ROCm-specific binaries by specifying when using . I had to extend version compatibility due to the specific requirements of ROCm 7.14, though there are no backwards-incompatible changes that kev requires that is no longer supported in the later pytorch versions.

TwelveNights · 1h ago
jaredpalmer
Round 26 result: no candidate (0 of 4); doubling tasksource-v1 left breadth and breadth ECE where round 25 had themMergedPR

Round 26's read-out, as registered in #181 (). The spec is unchanged. The full write-up, including the PLAN text, is in PLAN.md under "Round 26 result" and is committed in this PR. Verdict: no candidate (0 of 4). Nothing was confirmed, no test partition or locked set was read for round 26, and nothing is published. Every candidate fails . Breadth-v1 ECE is 0.0197 to 0.0287; the bar is 0.0176 (Kev-27B 0.0076). The final () passes 11 of 12 criteria and fails only breadth ECE (0.0233). Round 25's final did the same (0.0235). Breadth: +1.0 [+0.3, +1.7]. tasksource-heldout: +3.3. Kev panel: +5.6. Short states: −1.01, with a lower bound of −1.95 (inside the guard by 0.05 pp). CUAD: +0.2. The snapshots also fail the breadth primary (lower bounds −0.04 to −0.16 pp). s50 also fails short-state accuracy, and s25 also fails tasksource-heldout ECE. The single question was whether doubling tasksource-v1 turns round 25's small breadth gains into ones that pass. It did not (report only, ). Against round 25's lr 1e-6 arm at the same fraction of its run: Breadth is level at every point (−0.3 to +0.2, every interval spanning 0). Breadth ECE is level (Δ −0.004 to +0.000). tasksource-heldout is higher only at s25 (+1.7 [+0.7, +2.7]); at the final it is +0.5 [−0.5, +1.5]. The rank score is the same (+9.8 at the final). Short states lost on transfer-r3 test (final −1.2 [−2.2, −0.3]) and gained on transfer-v4 development (+1.2 [+0.2, +2.6]). Replay was 23.5 % against 30.2 %. Temperature, report only: Breadth still wants T 1.44-1.56, and tasksource-heldout wants 0.92-1.14. Unlike round 25, some temperatures inside the pooled 90 % interval meet both ECE bars for s50 and the final. For the final that is T 1.367-1.408, where the pool fit is 1.320. The rule serves the pool fit, so no verdict changes (). Against round 23's confirmed (report only; each side at its own pooled T; ): no arm is ahead with a lower bound above 0 on any gated panel. Against the final: Breadth −0.6 [−1.4, +0.2]. tasksource-heldout −0.9 [−2.2, +0.4]. Kev panel −2.6 [−3.6, −1.7], including hard-v1 −5.3. Short-state confident errors +0.9 [+0.4, +1.5] pp. Short states and CUAD accuracy are level. Training: one attempt on app , 8 × H200. The warm start loaded 850 full tensors plus the head from . The plan was exactly the projection: 458 steps, snapshots at 115 / 229 / 344, largest pass 25,590 of 40,960 padded tokens. No snapshot path needed correcting. Training took 5,836 s (1.62 h, against 1.77 h projected). Peak memory was 80.4 GB per GPU. Deviations (details in PLAN): The final's read batch went through the registered spend gate. The watcher launches a finished trial's reads without a gate, so it was stopped at 08:02Z while the trial trained. A helper polled the trial, ready to restart the watcher on a timeout, and restarted it at 10:13Z once s75 had launched and the gate passed. The gate counted whole batches at $400.97 each, as registered. Launch readings were $917.50, $1,346.60, $1,353.53 and $1,369.81 against the $1,373 limit (). Every read ran the first time. The private development partitions were in place before the deploy. Spend: Modal metered $4,558.69 at 13:01Z. That is +$104.88 over the $4,453.81 launch reading. The 11:17Z reading was $4,650.37 (+$196.56), and the meter later revised it down. The app is stopped, and was not touched. Private rows: , manifest . It holds the tasksource-heldout reads, the OOD rows, and the trial's development rows and full . No tasksource family name is in any committed file (grep-checked against the trial's task table, the tasksource-heldout rows' sources and tasks, and the exclusion list). Tests:** and pass locally (117 passed, 22 skipped). Do not merge without Jared's review.

jaredpalmer · 21h ago
jaredpalmer
Round 25 result in PLAN.md (the text #180 left out)MergedPR

PR #180 merged round 25's read-out files but not its PLAN.md write-up. This lands the "Round 25 result" section, the updated "Where we stand" bullet and Record row (no candidate, 0 of 8; lr 1e-6 s75/final fail only breadth-v1 ECE; replay held short states; none ahead of 27b-k-w85), and marks round 26 as launched. Generated with Devin

jaredpalmer · 1d ago
Structured data for AI agents

Repository: jaredpalmer/kev. Description: Jev-like family of decision models built on top of Qwen3.5/3.8 you can train and run on your own Stars: 7994, Forks: 504. Primary language: Python. Languages: Python (95.1%), TypeScript (4.3%), CSS (0.3%), Shell (0.2%), JavaScript (0.1%). License: Apache-2.0. Topics: decision-model, jev, qwen3. Latest release: kev-family (1w ago). Open PRs: 12, open issues: 17. Last activity: 21h ago. Community health: 42%. Top contributors: jaredpalmer, devin-ai-integration[bot], lunar-me, bhimrazy, TwelveNights, ImgBotApp, Radexito, VecSzn, kuishou68, mrcushen-arch.

·@ofershap

Replace github.com with gitshow.dev