A panel of four split on one candidate before Monday’s debrief
Marcus Bell gave this candidate a four for working with other teams. Joanna Reyes gave her a two for the same thing. A hiring manager can turn those two scorecards, and two more, into one page with Claude Cowork before the debrief, the meeting at ten where the panel decides. A scorecard is the form an interviewer fills in afterward. The form gives each competency, meaning one thing the job requires, a score from one to four, plus a few lines on what the interviewer saw. The candidate, the role and all four scorecards here are made up. The cards are wildly uneven. Renata Dolan wrote two hundred and seventy-one words of evidence. Tom Okafor wrote sixteen words, and two of Tom's entries read strong, four. A folder like this turns up after every round of interviews.
Anthropic’s interview kit ships its debrief template empty
| Role Competencies | “Define 4-6 key competencies for the role” · excerpt |
|---|---|
| Scorecard | “Rate each competency on a consistent scale (1-4) with clear descriptions of what each level looks like.” |
| Debrief Template | “Structured format for interviewers to share findings and make a decision.” |
| Output | “Produce a complete interview kit: panel assignment (who interviews for what), question bank by competency, scoring rubric, and debrief template.” |
Anthropic publishes a hiring skill for Claude called interview-prep. A skill is a written set of instructions Claude follows for one kind of job. The interview-prep skill writes the question bank, the plan for who interviews for what, a rubric that scores each competency from one to four, and a template for the debrief. Anthropic describes that template as a format for interviewers to share findings and make a decision. All of that gets used before anyone meets the candidate. Filling the template from four finished scorecards is still the hiring manager's job.
An average of four scores hides where the panel split
| Competency | Renata | Marcus | Joanna | Tom | If you averaged |
|---|---|---|---|---|---|
| Financial modelling | 4 | 4 | 4 | 4 | 4.0 |
| Stakeholder partnering | 3 | 4 | 2 | 3 | 3.0 |
| Systems and data | 3 | 3 | 3 | 3 | 3.0 |
| Communicating findings | 3 | 3 | 2 | 3 | 2.75 |
| Ownership through close | 4 | 4 | 4 | 4 | 4.0 |
With four scorecards open, the reflex is to average them. Four scores per competency become one number per competency, and the grid looks done. On this panel, that grid buries the argument the meeting exists to settle. Take stakeholder partnering, which here means how the candidate works with teams outside finance. The four interviewers scored it three, four, two and three. Now take systems and data. All four interviewers scored it three. Both of those competencies average exactly three point zero. On stakeholder partnering, Marcus Bell, a fellow analyst on the same team, gave a four, and Joanna Reyes, the director of operations, gave a two, about the same candidate on the same day. Three point zero cannot tell those two competencies apart. On a summary typed up by hand, the two rows sit next to each other and read the same.
The debrief page quotes each interviewer and opens on the widest split
“I asked twice how she would handle a plant manager who disputes her variance numbers in front of his own team, and both times the answer was to take it back to her director.”
“Walked me through renegotiating the forecast cadence with the commercial team after they missed two submissions”
So the job is extraction. For each competency, find the sentence each interviewer wrote about it, and carry that sentence onto the page word for word, with a name on it. The rule comes from Anthropic's own resume-screening guidance, which says: for each rubric criterion, find the line in the application that speaks to it, and quote it. A paraphrase of four people fails for the same reason an average does. Both hand you one voice where there were four. Under stakeholder partnering, Joanna Reyes asked twice how the candidate would handle a plant manager who disputes her numbers in front of his own team. In Joanna's words, both times the answer was to take it back to her director. Marcus Bell wrote about the commercial team, and in Marcus's words, she went to their standup every week until they stopped needing the reminder. Then order the page by spread, meaning the highest score any interviewer gave minus the lowest. Stakeholder partnering has a spread of two points, so stakeholder partnering opens the page. The three competencies where all four scores match sit at the bottom. The meeting starts on the argument instead of finding it on page three.
The Cowork project opens one folder holding four scorecards and the instructions
| Created from | an existing folder on this computer |
|---|---|
| Folder | debrief-alina-vos |
| Holds | 4 scorecards · PROJECT-INSTRUCTIONS.md |
| Approval | Manually approve |
“Projects you create from an existing folder on your computer stay on that computer and aren’t saved to your Claude account.”
The tool is Claude Cowork, Anthropic's workspace for handing Claude a piece of work, on the paid plans: Pro, Max, Team or Enterprise. Reading files on your own computer takes the Claude desktop app. There you make a project from a folder you already have, and Claude reads and writes the files in it with no upload step. Anthropic's help page says a project made that way stays on that computer and is not saved to your Claude account. Claude still processes the files on Anthropic's servers. In Manually approve mode, Claude pauses before each action, and you choose allow or deny.
The instructions file’s first rule forbids summarising any interviewer
| quoted passages that match the cards | |
|---|---|
| With the instructions | ✓ 20 of 20 |
| Without them | ✗ 1 of 4 |
The instructions are one page of plain English, saved next to the scorecards. One demand sits above the rest. Quote the interviewer's own sentence, and never summarise it. The page also rules out averages and any hire or no-hire line. I gave Claude these four scorecards twice, once with the instructions and once without. With them, all twenty quoted passages matched the cards word for word. Without them, one of four quoted strings matched. That draft put the words, take it to my director, inside quotation marks. Joanna Reyes's card says, take it back to her director. That is one test run of each, and that gap is why the quote demand leads the page.
The four interviewers settle the split in the room
Scorecard folder
- Renata Dolan
- Marcus Bell
- Joanna Reyes
- Tom Okafor
Quote each interviewer
- Financial modelling
- Stakeholder partnering
- Systems and data
- Communicating findings
- Ownership through close
Order by spread
Debrief page
Debrief meeting
Claude
locate the sentence each interviewer wrote
quote it word for word with their name
map it to the competency
order the page by spread
Hiring manager
read every quote back against the cards
resolve the split in the room
decide whether to raise any off-work line
own the hire/no-hire call
Here is where I'd stop and read it myself. The handoff line marks where the work passes from Claude to you. On Claude's side, Claude finds each interviewer's sentence, quotes it with a name attached, files it under a competency and orders the page. On the hiring manager's side, the first job is reading every quote back against the cards. Then the panel works out whose reading of stakeholder partnering holds up, Marcus Bell's or Joanna Reyes's, and the hiring manager owns the hire or no-hire call. Keeping that call off the page is a rule in the instructions. The test draft without instructions ended on three options: hire, a follow-up conversation first, or no hire.
A “one of us” remark stays on the scorecard, off the debrief page
“Felt like one of us from minute one”
“Would be glad to have her at the Thursday lunch.”
“be cautious about granting access to sensitive information like financial documents, credentials, or personal records.”
“Consider creating a dedicated working folder for Claude rather than granting broad access”
Two parts of this job belong to you. The first is any line about something other than the work. Tom Okafor wrote that the candidate felt like one of us from minute one. Anthropic's own screening guidance names the phrase, seems like one of us, as a place where unexamined preference hides. The instructions leave that kind of line out, and the hiring manager decides whether to raise it. The second part is the folder. Anthropic's safety page says to be careful giving Claude access to personal records, and suggests a dedicated working folder. Four people's evaluations of one named candidate are personal records, so this folder holds the four scorecards, the instructions, and nothing else.
By hand the time goes to retyping, and with Cowork it goes to reading
Both numbers are estimates, itemised against these four scorecards. Done manually, the job takes about thirty-five minutes, and about twenty of those minutes go to finding and retyping twenty passages. With Cowork drafting, the job takes about twelve minutes. About eight of those minutes go to the read-back, where you check each of the twenty quotes against its card and confirm the page carries no average or verdict. Checking the quotes is the job, so the check stays on the timesheet.
Monday’s debrief opens on stakeholder partnering, with both sentences on the page
“I asked twice how she would handle a plant manager who disputes her variance numbers in front of his own team, and both times the answer was to take it back to her director.”
“Walked me through renegotiating the forecast cadence with the commercial team after they missed two submissions”
At ten on Monday, the hiring manager walks in with Marcus Bell's four and Joanna Reyes's two at the top of one page, with each interviewer's own words under each score. The meeting goes to how this candidate would handle a plant manager who pushes back in front of his own team. The four people in the room make the hire call.













