Paste a job description into an AI interview question generator and a full list of questions comes back almost immediately. They read fine. They also produce screens where two interviewers hear the same answer and score it differently, because nobody wrote down what a good answer contains. This walks through a workflow that turns a job description into a small interview kit a hiring manager will approve and a recruiter can score the same way twice.
A list of questions is the wrong output
The question list feels like the deliverable because it's the part that's easy to generate. It isn't the part that makes an interview work. What decides that is the pairing of a question with a fixed follow-up probe and a written description of what a 1, a 3, and a 5 actually sound like.
The U.S. Office of Personnel Management's structured interview guide sequences it the same way: analyze the job, pick competencies, then write questions and probes, then build rating scales, then pilot the thing before anyone uses it on a real candidate. The questions sit in the middle of that chain, not at the start of it.
So the useful output looks like this, per competency:
job description -> evidence map -> question -> fixed probe -> 1/3/5 anchors -> pilot
Five rows of that is a working interview kit. A long list of unanchored questions is a document nobody opens twice.
Step 1: Pull 4 to 6 interviewable competencies from the job description
Most job descriptions carry a lot of boilerplate. Strip the culture paragraph, the benefits list, and the "other duties as assigned" line. What's left is a much shorter set of real requirements, and only some of them belong in an interview at all.
Sort each requirement by where the evidence actually lives:
| Requirement type | Where to check it | First screen? |
|---|---|---|
| Hard gate (license, shift availability, location) | Application field or one yes/no question | Yes, as a filter, not a scored question |
| Tool or platform experience | Resume plus one confirming probe | Only if it's a genuine gate |
| Craft skill (writing, forecasting, debugging) | Work sample or later panel | No |
| Recurring behavior (conflict, prioritization, ownership) | Structured interview question | Yes |
| Claimed results and track record | Reference check | No |
What survives that sort is your competency list. Cap it at four to six. If a hiring manager insists on nine, ask which two they'd drop if the screen had to fit in 25 minutes; they'll usually cut three.
One common failure here: asking one question per job description bullet. Twelve bullets becomes a twelve-question screen, which becomes a 40-minute call that a recruiter running six screens a day cannot sustain. Four scored questions plus two filters is a real first round. Seven scored questions is a hiring manager interview wearing a screen's name tag.
Step 2: Define what good looks like before the AI writes anything
This is the step teams skip, and it's the reason generated questions come out generic. A vague job description produces vague questions because the model has nothing else to work from.
Sit with the hiring manager for ten minutes and get concrete answers to three things per competency: what a strong person does in the first 90 days, what a weak hire gets wrong, and what a specific recent example of good performance looked like on their team.
For a warehouse shift supervisor, that conversation produces material like this. Strong supervisors reassign labor in the first minutes of a short-staffed shift instead of waiting to see how it goes. Weak ones let outbound slip and tell the next shift after the fact. A real example from last quarter: someone pushed restock to the following shift, protected the outbound cutoff, and flagged it at handover.
That's an evidence map. Now the model has something to work with, and the questions stop sounding like they came from a template. It also gives you material for anchored scorecards that mean the same thing to every interviewer.
The prompt template
Copy this, fill the three input blocks, and run it. The constraints matter more than the phrasing.
You are helping build a structured first-round interview for one role.
INPUTS 1. Job description: [paste full text] 2. Hiring manager evidence map: [for each competency, what a strong answer contains, what a weak hire gets wrong, one real example] 3. Seniority and screen length: [e.g. frontline supervisor, 25 minutes]
RULES - Produce at most 5-7 questions total for a first screen. Fewer is fine. - Every question must map to one competency from my evidence map. Do not invent competencies or requirements that are not in my inputs. - If the job description is too vague to write a scoreable question for a competency, do not guess. List it under MISSING CONTEXT and tell me the exact question to ask the hiring manager. - Do not ask about personal circumstances that are unrelated to the job, and do not ask about characteristics protected under applicable law (for example age, family or parental status, pregnancy, national origin, religion, disability, or health). Do not ask proxy questions that surface them either: graduation year, "where are you originally from", childcare plans. Treat this as a floor, not a complete list. - Do not include criminal-history or arrest-record questions. Rules differ by jurisdiction, and any such question needs legal review and a documented job-related reason before it goes in a screen. - Behavioral or situational only. No trivia, no puzzles, no questions answerable from the resume alone. - Calibrate to the stated seniority. An IC-level question asked of a manager wastes the slot. - Each question gets exactly one fixed follow-up probe that everyone asks.
OUTPUT A table with these columns: Competency | Question | Probe | 1/3/5 anchors Anchors must describe observable behavior in the answer, not adjectives. "Names the tradeoff they accepted" is an anchor. "Shows strong judgment" is not.
Then a MISSING CONTEXT list.
The two clauses doing the heavy lifting are the missing-context rule and the anchor definition. Without the first, the model fills gaps with plausible invented requirements. Without the second, you get anchors like "excellent communication" at 5 and "poor communication" at 1, which is the same unscoreable interview you started with.
The constraint list keeps obvious problems out of a draft. It is not legal coverage, and it is not a substitute for having your own counsel confirm what's allowed in the places you hire.
Review every generated question: pass, clarify, stop
Run each question through this before it goes near a candidate. It takes seconds per question.
| Outcome | Trigger | Action |
|---|---|---|
| Pass | Job-related, asks for observable evidence, matches the seniority, anchors are distinct | Keep as written |
| Clarify | Reasonable question, but the role context or the manager's standard is missing | Send to the hiring manager with a specific question |
| Stop | Touches a protected characteristic or unrelated personal life, invents a requirement absent from the job description, duplicates another question, or can't be scored consistently | Delete, don't rewrite |
The invented-requirement case is worth watching for specifically. Ask for warehouse supervisor questions and models will confidently produce something about forklift certification and OSHA recordkeeping, whether or not either appears in your job description. If it isn't a real requirement, testing it screens people out for nothing.
The EEOC's guidance on employment tests and selection procedures is blunt about where responsibility sits: the employer is on the hook for job-relatedness and validity even when an outside vendor supplied the tool and the documentation. An AI interview question generator does not transfer that to the model.
Bad versus good, on a real question
Take the shift supervisor role and the competency "handles labor shortfalls without losing the outbound cutoff."
Bad version, which is roughly what you get from a raw job description with no evidence map:
"Tell me about your leadership style and how you motivate your team."
It's unscoreable. Every candidate says some blend of leading by example and open communication. There's no observable behavior in the answer, so the interviewer scores tone of voice and calls it culture fit.
Good version:
"Walk me through the last shift that started short-staffed. What did you change in the first hour, and what happened to your numbers that night?"
>
Fixed probe: "What did you decide to let slip, and who did you tell?"
Anchors:
- 1: Describes the situation only. Actions belong to "the team" or "we." No decision they personally made, no outcome.
- 3: One concrete action (moved two pickers from restock to outbound) but no numbers and no reasoning about what they gave up.
- 5: Action, the tradeoff they accepted, the result with a number they own, and who they told. For example: held the outbound cutoff at 92% of plan, pushed 40 pallets of restock to the next shift, flagged it to that supervisor at handover.
Same competency, same 90 seconds of candidate airtime. One version can be scored the same way by two different interviewers; the other can't. That's the difference between a question list and a working interview scoring rubric.
Pilot it for 20 minutes before it goes live
Book 20 minutes with the recruiter and hiring manager and do four things.
Read every question out loud. OPM's guidance on creating structured interview questions recommends exactly this, and it catches more problems than silent review does. Questions that scan fine on a page turn out to be two questions stapled together, or full of internal jargon a candidate outside the company won't parse.
Answer the questions yourselves. If the hiring manager can't produce a 5-level answer to their own question, the anchors are wrong or the question is.
Cut duplicates. Generated sets routinely contain two questions probing the same competency from slightly different angles. Keep the one with better anchors.
Time it. Multiply questions by four minutes and add five for intro and candidate questions. If that exceeds your screen length, cut a question rather than rushing the probes, because the probes are where the scoreable detail comes from.
Then name an owner. One person approves the final kit and one person owns the hiring decision. The model drafts; it doesn't sign off, and it doesn't decide.
What an AI interview question generator is genuinely good at
It's fast at producing a broad, role-relevant first draft, which beats staring at a blank document. It's good at turning a manager's messy verbal standard into anchor language. It's good at spotting that you have four questions all testing the same competency.
Where it's weak is predictable. Vague inputs produce vague output, every time. Seniority gets mislabeled, and a set written for a "manager" comes back full of individual-contributor questions because the job description reads like an IC posting with a manager title on top. Technical questions drift toward trivia that tests recall rather than how someone does the work. And a well-written question with generic anchors still scores inconsistently, which is the failure most teams never notice because the questions look so good.
Once a kit is approved and calibrated, it's stable enough to run the same way for every candidate in the pipeline, which is the point where AI-assisted candidate screening becomes useful: the same questions, the same probes, the same anchors, whether the candidate applied Monday morning or Friday at 11pm. Fitting that into the wider workflow is covered in this guide to structured interview questions.
Revisit the kit when the job changes. A question set built for a role that has since absorbed two new responsibilities is measuring a job that no longer exists.
Key Takeaways
- The output worth generating is a compact kit, not a question list: competency, question, fixed probe, and 1/3/5 anchors written in observable behavior.
- Extract four to six interviewable competencies first, and route entry requirements, craft skills, and track-record claims to forms, work samples, and reference checks instead of the screen.
- Get the hiring manager's evidence map before prompting. Vague inputs are the single biggest cause of generic AI-generated interview questions.
- Force the model to flag missing role context rather than invent requirements, cap it at five to seven first-screen questions, and demand anchors that describe behavior, not adjectives.
- Run every question through pass, clarify, or stop, and delete anything that touches protected characteristics, invents a requirement, duplicates another question, or can't be scored consistently. Anything involving criminal history goes to legal review before it goes in a screen.
- Pilot for 20 minutes: read aloud, answer your own questions, cut duplicates, check timing, and name the human who approves the kit and owns the decision.
