Another week, another pile of repositories crossing the ten-thousand-star mark in days. Most weekly roundups stop at the star count, but the interesting story this week is what the stars are pointing at: a cluster of projects built around a simple idea — that most of what an LLM does in an agentic loop is make small, structured decisions, and that those decisions do not need a text-generation model at all.
Here are five repositories that trended hard this week, what they actually do, and why the pattern behind them matters more than any single project.
1. Laya — Typed Decisions in One Forward Pass
Laya is a non-autoregressive decision engine: instead of generating text and parsing an answer out of it, it answers typed questions — a choice among labels, a score, or a yes/no probability — in a single forward pass over the input. Around 33 milliseconds per question on a T4, roughly 7 milliseconds per question when batched. The base checkpoints are compact encoder models (a 421M-parameter ModernBERT-large for English, a 322M mmBERT-base covering over 100 languages), and a router inspects each request and picks the right checkpoint.
The training recipe is the interesting part. The checkpoints are trained with reinforcement learning against strictly proper scoring rules, which means the model is rewarded for outputting calibrated probabilities rather than confidently wrong ones. The API makes that calibration a first-class output:
from laya import Router
router = Router()
state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
"department": {
"type": "choice",
"instructions": "Which department should handle this?",
"criteria": {
"billing": "invoices, payments, refunds",
"technical": "bugs, outages, system errors",
"other": "everything else",
},
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["not urgent", "soon", "blocking"],
},
"churn_risk": {
"type": "noul",
"instructions": "Does the user threaten to cancel or leave?",
},
}
result = router.predict(state, questions)
print(result["answers"]["department"]["choice"]) # billing
print(result["answers"]["churn_risk"]["noul"]) # probability the answer is yes
Because there is no text generation, there is nothing to parse and nothing to hallucinate — the output is either a valid label from your set or a probability. The project also ships an opt-in abstention mechanism (min_confidence= flags low-confidence answers instead of guessing) and a fine-tuning notebook that runs on Kaggle’s free T4s. On the project’s typed-decisions benchmark, fine-tuning lifts accuracy from 0.362 to 0.766 — a reminder that the zero-shot numbers are a floor, not the product. Apache-2.0 licensed, with an MLX companion runtime (laya-mlx) that pushes short decisions down to the 7–14 ms range on Apple silicon.
2. Jev Ultrafast — A Browser Agent With an Indexed Action Space
The Browser Use team’s demo made the rounds for the headline number: a natural-language flight search on Google Flights completing in about seven seconds, loading waits included. The mechanism is more interesting than the number. Instead of showing an LLM a screenshot and asking “what should we click?”, the agent converts every page observation into a numbered element table and lets a small decision model pick an operation (CLICK, TYPE_TEXT, SELECT, SCROLL_UP/SCROLL_DOWN, WAIT, DONE, BLOCKED) and a target — and only falls back to a text-generation model when the chosen operation actually requires typing text.
Operation and target are decided in one network round trip, because both heads see the same observed state. The default loop never touches screenshots at all: the model consumes structured state, and the geometry of the selected target is validated against the live DOM before input — animated or occluded controls get rejected rather than triggering another slow prediction cycle. The Python API is a handful of lines:
from jev_ultrafast import Agent
with Agent(
"https://www.google.com/travel/flights?hl=en",
"Find one-way flights from Zurich to London on September 20, 2026, "
"for one adult in economy. Stop when matching flight options are visible.",
) as agent:
for state in agent.run():
print(state["elapsed_ms"], state["status"])
This is the same philosophy as Laya applied to browser automation: reserve expensive token generation for the one step that needs it, and make everything else a fast structured decision over an indexed action space. MIT licensed.
3. ZCode — Z.ai’s Open Coding Agent Harness
Z.ai — the company behind the GLM model family — open-sourced its coding agent workspace this week, and it cleared 6,000 stars in a few days. ZCode is a full harness: an Electron desktop app, a browser UI, a terminal TUI, and the agent CLI and runtime itself, all in one TypeScript monorepo under an Apache-2.0 license.
One binary drives all three interfaces: zcode with no arguments opens the TUI, zcode --web starts a local web workspace on port 3030, and the remaining arguments go to the agent CLI. The repo is dated and fast-moving — v3.14.3 landed within days of the initial release — and the README is notably operational about its toolchain, pinning Node.js and pnpm versions through a mise.toml file rather than prose instructions.
For anyone building on GLM models, this is the vendor’s own answer to the “which harness do I use” question, with the client, backend, and runtime source all inspectable rather than distributed as an opaque binary. It is also a data point on where the coding-agent market is heading: the harness layer is becoming open infrastructure, and differentiation moves to the models underneath.
4. Fast Jev Compaction — Scoring Context Instead of Summarizing It
Coding agents have a context-management problem: long sessions hit the window limit, and the standard fix — asking the model to summarize the conversation so far — is slow, lossy, and throws away exactly the operational detail (tool outputs, file states) that the agent will need next. This Claude Code plugin replaces the compaction summary with a typed-decision pass: every tool call and result in the session is scored in a single fast request, stale entries get dropped, and what survives is a structured decision record rather than a prose paraphrase.
It is a small project (MIT licensed, TypeScript) but it points at something bigger. Summarization is a compression format optimized for human readability; decision records are a compression format optimized for the next model turn. Expect more tooling in this niche — context compaction is now a hot path in every long-running agent, and hot paths attract engineering.
5. Reladraw — Diagrams Where Layout Is a Choice
A palette cleanser to finish. Reladraw, at 389 points on Hacker News this week, is a diagram language built on one contrarian bet: automatic graph layout is usually wrong, and the person drawing the diagram should decide where things go. Text-based diagram tools are either fully declarative (Mermaid — the tool positions your boxes, often badly) or fully manual (draw.io — pixel dragging). Reladraw keeps the text-file source of truth but makes position a first-class part of the language, so a diagram’s structure and its layout live in the same diffable file.
That matters for exactly the workflows where diagrams hurt most today: architecture decision records and system-design docs in version control, where an auto-layout engine re-scrambling your carefully arranged boxes on every edit is a real, recurring annoyance.
The Pattern: Small Models for Structured Decisions
Step back from the individual projects and one theme ties four of these five together. The current agent stacks spend frontier-model tokens on decisions — route this ticket, click that element, keep or drop this context entry — that are structurally classification problems. This week’s trend data says the ecosystem is noticing. Typed decision heads, indexed action spaces, and calibrated probabilities are cheaper, faster, and in several of these projects more accurate than asking a large autoregressive model to emit its answer as prose and hoping the parse succeeds.
None of this replaces text generation where it counts — reasoning, writing, synthesis. It slices the agentic loop in two: generation stays with the big model, and the hundreds of small routing decisions around it move to models measured in hundreds of millions of parameters and tens of milliseconds. That split is probably the most practical architecture idea to come out of this week’s trending page.