Open source

GitHub Issues Summarizer

Paste a repo URL, get the briefing its issue tracker never wrote

Try it liveView the repo
  • Next.js 16
  • TypeScript
  • OpenAI
  • Cloudflare Workers
  • Workers KV
  • GitHub GraphQL
  • Tailwind v4

A word on what this is before anything else: a small exercise. One simple workflow with a handful of nodes on it, published as a demo you can run and pull apart. It is not a large system and it is not trying to be. Small is what makes the moving parts legible.

Open an unfamiliar repository’s issue tracker and you get a hundred threads and no picture. Paste that repo’s URL into this and you get one page instead: the dominant themes, what needs attention, the questions nobody has answered, and a maintenance score out of 100 for the repository. Every claim is anchored to a real issue number, and every number links back to its source.

The summarising is the least interesting part of it. What this is really for is the control flow around the model calls: where the work branches, what runs in parallel, what checks the output before a human ever sees it, and what stops the whole thing from spending money it does not need to spend.

It is also deliberately not an agent. It is a workflow with nodes, and I want to be precise about that, because the word agent is doing a lot of unearned work right now. The path here is fixed at design time. No model decides what to do next; models are called at four fixed points to do bounded tasks. When you can enumerate the steps in advance, that is the cheaper and more predictable choice, and a good number of the systems being sold as agentic would be better off built this way. Use the least autonomy that solves the problem.

4
model calls in the path
100
issues per run, ceiling
2
passes, then it accepts
0
calls on an unchanged repo
67%
cheaper after one model swap
102
tests, no credentials needed

The architecture

Repo URL

Any public GitHub repository. The visitor never authenticates.

Routerno model

Have I summarised this repo before? A key lookup, not a judgement call.

Scoutno model

GraphQL pulls up to 100 open issues and 20 comments each in two requests. REST would need about 101.

Updaterno model

Compares every issue’s updatedAt against the cached copy and re-fetches only what moved.

Parallel fan-out

LLM 1 · titles and bodiesgpt-4o-mini

Mechanical extraction. High volume, low judgement.

LLM 2 · comment threadsgpt-4o-mini

Same profile, run concurrently. Latency is the slower branch, not the sum.

LLM 3 · coherence checkergpt-4o-mini

Reads the digests back against the source and flags anything invented or incoherent.

Flagged issues go back to 1 and 2: those only, at most twice

Health scorerno model

Turns repo-wide counts and the sampled issues into a maintenance score out of 100. A pure function, so the same repo always scores the same.

LLM 4 · composergpt-4o

Themes, what needs attention, open questions, and the score explained. The part a person actually reads.

Score and briefing

Streamed to the browser stage by stage, rather than a two-minute spinner.

KV cache

Held for seven days, which is what makes the next run cost nothing.

One fixed path, ten nodes, four of them model calls. Two are named orchestration patterns doing real control-flow work; the rest are plain steps, and calling any of them an agent would be a lie.

How a run actually goes

The router classifies the request into one of two paths by checking the cache. Cold means I have never seen this repository, so the Scout fetches the tracker. Warm means I have, so the updater compares timestamps and re-fetches only the issues that moved, dropping the digests of any that have since closed.

From there the two extraction stages run at the same time (one on titles and bodies, one on comment threads) because neither needs the other’s output. Their digests go to the coherence checker, which is the only place in the system with genuine iteration: if it flags issues, those issues, and only those, are summarised again. The loop is capped at two passes, after which the best digest available is accepted. An unbounded fix-until-perfect loop is an infinite loop and an unbounded bill wearing a nicer name.

Between the loop and the composer sits the one step with no model in it at all: the scorer. Then the composer writes the briefing from verified digests, the result streams to the browser as it is produced, and a copy goes to the cache on the way out.

The score is computed, not opined

The page opens with a maintenance score out of 100, and no model produces it. It is a pure function over repo-wide counts and timestamps, so the same repository always scores the same and every point traces back to a named component with a stated basis. The model receives the score and its breakdown and explains them. It does not derive them, and it cannot move them.

That split is the whole idea. Asking a model to rate something out of 100 gets you a number that sounds considered and moves between runs for no reason you can name. Compute it, and you can be asked to defend it.

The components are ratios and recency, never raw counts, because counting open issues measures adoption rather than health. Two real results make the point: mitt has 16 open issues against 99 closed, nothing closed in the last 90 days and no push in two years, and scores 47. Zod has 56 open against 3,096 closed, 117 closed in the last 90 days and a push the same day, and scores 97. By issue count mitt looks like the healthier project. It is dormant.

The honest limit ships with the score: it reads the tracker, so it measures maintenance, not code. A quiet project with excellent code scores low, which is the right answer to whether anyone is looking after it and the wrong answer to whether it is well written. The score says so itself.

The decisions worth arguing about

GraphQL instead of REST for the fetch. REST needs roughly 101 requests per repository: one for the issue list, one per issue for its comments. GraphQL brings issues and comments back together in about two. The consequence is that the API refuses unauthenticated calls, which is why the operator supplies a token and the visitor never has to.

One model everywhere is a habit, not a decision. Extraction is high-volume and mechanical, so it runs on the small cheap model. The composer, the one output a human reads, runs on the strong one. Every stage is assigned in a single file that also carries each model’s price, so a run can report what it cost.

The verifier used to run on the strong model too. I planted defects in the digests and measured whether the cheap model caught them: it matched the expensive one, at a seventeenth of the price. Moving it cut a cold run on Zod from ten cents to three, a 67% saving, and the smaller repositories fell 35% and 41%. Two things fell out of that exercise that I did not expect. The extraction stages were never the cost driver, they are 10% to 30% of a run. And newer or smaller does not mean cheaper: one nano model spends roughly 1,850 hidden reasoning tokens on a 180-token extraction and bills them as output, and another was rejected outright after catching 0 of 4 planted defects as a verifier and silently dropping issues as an extractor.

The coherence checker derives its verdict from the list of problems it enumerated, not from the boolean it reports about itself. It is a small thing, and it matters: never let the evaluator grade itself on a flag it can flip.

Cost is a design constraint, not an afterthought

Every cold run spends real money, and the page is public, so the spend is bounded in four places rather than hoped about in none. Per request: at most 100 issues, at most 20 comments each, bodies truncated at 4,000 characters and comments at 1,500. Per visitor: a per-IP quota on cold runs, with cache hits deliberately free so an already-summarised repository keeps working even past the limit. Per repository: the change-detection cache, which is the real saver, because traffic to a portfolio piece hits the same few repos over and over. And behind all of it, a hard spend ceiling on the account, which is the only control that does not depend on my own code being right.

A run over 100 issues is roughly 15 to 25 model calls and takes one to three minutes, almost all of it waiting on I/O. A repeat run on a repository nobody has touched is zero model calls and arrives immediately.

What the output is allowed to do

The briefing has a contract, and each rule in it is enforced at the level it can actually be enforced. Mechanical rules run in code after the model returns. Rules that need judgement stay in the prompt, where they belong.

An issue may appear under exactly one theme, so a deduplication pass enforces it: asking nicely got the partition right about two thirds of the time across three repositories, which is not a rule, it is a tendency. Only issue numbers that appear in the verified digests become links, so an invented number cannot turn into a plausible-looking source link. The link itself is derived from the repo and the number at compose time, so nothing extra is stored and no tokens are spent producing URLs.

The prose rule I keep on this site is in there too: the composer’s output has its em dashes stripped in code, because telling a model not to write one is not the same as it not writing one.

One rule resists both levels. The composer still names a theme after severity now and then, even handed a clean subject vocabulary and told not to. Those are detected and logged rather than rewritten, because choosing a replacement name is a judgement call, and renaming a heading the surrounding prose refers to would trade one incoherence for another. Logging what you could not fix is more useful than quietly papering over it.

What I am honest about in the README

The endpoint sits behind a same-origin check, a short-lived signed token minted by the page, and a per-IP quota. The token is a speed bump, not a lock, and the README says so: this is a public page with no login, so the browser has to be handed a token, and a person can open devtools, read it out and replay it until it expires. It stops bots, scrapers, direct callers and anyone embedding the endpoint from another site. The quota is what actually bounds the bill. Closing the last gap means tying the token to an identity, which means an identity provider, which is documented as the upgrade path rather than pretended away.

Issue text is untrusted input. It comes from arbitrary strangers, so an issue titled with a script tag can reach the summary; the renderer escapes raw HTML, restricts link schemes and renders images as alt text only. Treating model output as untrusted is a security decision, not a formatting one.

And failures are recorded rather than hidden. One batch falling over does not abort the run: those issues get marked as failed, are excluded from verification and from the briefing, and the count is reported. Only a total failure, which is systemic, stops everything.

The design document came first

The architecture document in the repository was written before any code existed, then handed to the coding agent to implement, and updated whenever a decision changed during the build. The model named in the original design was retired mid-project, so the document changed, not the other way around.

That is the part I would point at if you only have time for one thing. The agents are fast and tireless and genuinely good at their pieces. Deciding what the system is, where it may branch, what bounds it, and what counts as correct is still the work, and it is still mine.

Try it liveView the repo

Built something like this, or want one built? Send me a message >

All code