Engineering managers run performance reviews under conflicting pressures:
- Remember a year of work for 6–12 people in 2–3 weeks
- Write fair, specific, growth-oriented feedback without legal/HR landmines
- Compare people to published role expectations and the next progression
- Connect reviews to promotion discussions — without those being the same form
- Do it while still shipping product — often without waiting on IT for a new SaaS tool
Existing tools fail EMs in predictable ways:
| Tool pattern | Failure mode for EMs |
|---|---|
| HRIS annual form | Blank boxes, no engineering context, no evidence trail |
| Docs + spreadsheets | No process, no queryable history, calibration chaos |
| Continuous feedback apps | Lots of kudos, weak cycle structure and rating rigor |
| Generic AI writers | Polished empty prose; invents achievements; hides bias |
The best tool for an EM is a local workbench: a living engineering dossier that becomes a structured judgment packet at cycle time, judged against uploaded role & responsibilities, with AI that argues from evidence.
v1 scope lock: solo EM workbench. Multi-party trust features (live anonymity, upward aggregation, org HR audit, adverse impact) are hosted-only later. See TRUST_MODEL.md.
Owns 4–12 direct reports. Uses the desktop app daily/weekly. Needs high-quality reviews fast, promo cases grounded in next-role expectations, and continuity across cycles.
Completes self or peer forms via exported bundles; reads a shared packet the EM exports/sends.
Cross-manager calibration and org policy — hosted or bundle-calib, not assumed on one laptop.
- Before the cycle: Capture evidence continuously; backfill history on first install.
- At kickoff: Open a cycle, pin role assignments, export self/peer bundles.
- During writing: Draft manager reviews from evidence + imports in ≤45 minutes/person.
- Progression: Compare to next RoleDefinition; open a promotion packet when warranted.
- At share-out: Export a clear packet for the IC discussion.
- Afterward: Query history on this workspace — trends, stagnation, prior themes.
- Evidence over eloquence. Prefer linked artifacts to adjectives.
- Role docs are the bar. Uploaded R&R define current-role fit and next-progression comparison.
- Performance ≠ promotion. Linked via the same framework, not one score.
- Honest about trust. Do not promise anonymity or HR controls the local file cannot enforce.
- AI is a co-pilot. Humans own final text and ratings; cloud egress is opt-in per data class.
- Manager time is scarce. Optimize EM throughput; import paths over multi-user ceremony in v1.
- History compounds. Every cycle should make the next easier — including cold-start backfill.
- A generic survey tool with a “performance” skin
- An AI that writes entire reviews from a job title
- A forced-ranking machine
- A surveillance product that scrapes private Slack DMs
- A fake multi-user HRIS that stores “anonymous” data in a file the manager owns
- An HRIS/payroll replacement
| Metric | Target | How measured (local) |
|---|---|---|
| Median manager time per finalized review | ≤ 45 minutes | In-app writing-desk timer (on-device only) |
| % non-middle ratings with ≥1 evidence link | ≥ 80% | Local workspace stats |
| Bundle import success for self reviews | ≥ 90% of directs | Local |
| % ratings judged with assigned RoleDefinition | 100% (warn if unassigned) | Local |
| AI draft used then edited (when AI on) | 40–70% accept-with-edit | Local |
| Design-partner qualitative NPS | Track in pilot surveys | External pilot, not product telemetry |
Product-wide NPS/telemetry is not assumed from the local app unless a future opt-in exists.