"Performance" is a loaded word in software. Say it out loud and most people hear a review, a rating, someone above you deciding how good you are. We measure performance, but not yours. We measure the work's. That distinction is the whole stance, so it is worth being blunt about what we do and do not do. What gets attributed, and to what Seorak attributes outcomes to the work, not to you. The unit is a session and what came out of it. Did the run reach a commit. Were the lines it changed still on the branch later, or had they been replaced. How long it ran, what it cost in tokens, where it stalled. None of those say anything about you as a person. A session that churned might be the right call on a hard problem. A session that survived might have been a rename. The stats describe what happened to the code, and that is all they claim to describe. We measure the performance of the work, not a review of the person. There is, deliberately, no score. No number that grades you, no weekly figure that goes up or down to tell you how you are doing. We considered it. A single score is easy to build and easy to read, and it is exactly the thing that turns honest measurements into a verdict. We did not invent that worry. When a measure becomes a target, it stops being a good measure, because effort moves to the number and away from the thing it stood for. So there is no score field. There is a record, and you read across it, because the work was never one dimension to begin with and does not reduce cleanly to a single figure. When there is nothing honest to say The harder discipline is the empty case. Most attention tooling fills space. If there is not enough data, it shows a placeholder, a flat line, a zero that looks like a result. A read over time is a distribution, and a distribution needs enough observations to mean anything. So when there is not enough to say, the read stays empty and says so. A handful of sessions on a new project will not produce a confident statement about when you work well with an agent. The honest answer there is "not yet," and we would rather show that than invent a trend. This matters more than it sounds. The moment a tool starts asserting patterns it cannot support, it stops being a record and starts being an oracle. We would rather under-claim. Here is the line we hold: We attribute to, We do not produce What shipped (reached a commit), A score for the developer What survived (lines kept on the branch), A grade of you as a developer Where a session stalled or churned, A pattern when the data is thin Teams, without the manager seat This is where "performance" usually goes wrong, so we drew the line early. For a team, the read is self-scoped. You see your own work over time. Identifiers are salted hashes, opt-in, scoped to you. There is no manager view that ranks people. We did not forget to build it. It refuses to. The product would be more sellable to a certain kind of buyer if it leaderboarded a team, and that is precisely the product we are not making. Across field after field, what gets measured and rewarded gets gamed, usually at the expense of the thing the number was meant to protect. A leaderboard would be the shortest path to it. The buyer here is the developer, and the developer is the only person the read is for. What you get out of that is narrow on purpose: An honest record of a session, including where it stalled and whether the change held. A read over time of which conditions tend to produce work that lasts, framed as a distribution, never a grade. Nothing about anyone else's sessions, and nothing for anyone above you to read about yours. The point The reason to keep "performance" in the vocabulary at all is that the work does have performance. Some sessions ship and hold. Some run for an hour past the point they were helping. That is real, and worth seeing plainly. The trick is to keep the word pointed at the work and never let it drift onto the person. A commit survived or it did not. That is a fact about the branch. It was never a fact about you.