Back

Phenom

Faster candidate evaluations with AI Insights

Applied AI, built to give evaluators candidate insights for faster, more informed decisions.

CompanyPhenom
Year2026
Expertise
Applied AIAI PrototypingSystems thinking
Industry
HR TechEnterprise SaaSWeb app
Team

John SmithNDA Protected Product Manager

Laura SmithNDA Protected UX Designer

Adrian Burlău Product Designer

Disclaimer: This project is under NDA, the screens are recreations for portfolio purposes not the shipped product, but reflective of the actual design solutions I developed.

01 | Overview

Overwhelmed evaluators, rushed calls

At Phenom, I helped shape an AI-native feature that reads a candidate's screening interview and gives evaluators structured insights to work from, without replacing their own judgment. I owned the working prototype and AI-output testing, using four recruiter personas to evaluate how different AI approaches behaved before we committed to a direction.

02 | Problem

Fair evaluations, just not enough time for all of them

Manual scorecards worked at low volume. At hundreds of candidates per role, evaluators didn't have time to give every answer a fair, in-depth read and judgment quality varied day to day.

57%

Of recruiters and hiring managers struggle with candidate screening and assessments

48%

Of hiring managers attribute bad hires to the pressure of filling a position quickly

Evaluate candidates modal showing per-question 1-5 star ratings with no AI Insights tab
Manual evaluation

Hundreds of candidates, not enough time

Questions & answers only, no AI Insights tab. Per-question 1-5 star ratings, an aggregate "average rating" up top, nothing else to go on.

03 | Opportunity

Assist judgment, don't replace it

Customers were asking for faster decisions at scale, but the harder product problem sat underneath: how much authority should AI have in a hiring decision?

We explored whether AI should verdict the evaluation, sit passively as another source of information, or act as an assistant that gives recruiters useful evidence while keeping the final judgment with them.

The direction I advocated for was the third: AI should make evaluation faster and more informed without becoming the decision-maker.

How do you give evaluators a genuinely useful AI analysis, without teaching them to stop applying their own judgment?

04 | Solution

Confidence in the evaluation, not a final verdict

Lock 3 of a job's 10 attributes for the AI to evaluate while configuring the screening. Every completed screening opens on AI Insights by default, showing its confidence, never a yes/no of its own.

Screening configuration showing a questionnaire and 3 locked evaluation criteria out of 10 job attributes
Configure screening, once per job

3 of 10 attributes, locked

Standard or Conversational screening mode, a questionnaire, and 3 of the job's 10 attributes locked in for the AI to evaluate, set once, applied to every candidate for that role.

Candidate summary card followed by Advancement strengths and Development areas analysis
Candidate summary + Analysis

Summary before verdict

Who this candidate is, in plain terms, then the candidate analysis: Advancement strengths and Development areas, 3-6 bullets each. Context before anything recommendation-shaped, to reduce anchoring.

Core skills assessment showing Low, Medium, and High confidence per locked skill with expandable reasoning
Core skills assessment

Confidence in the read, not a score

Low / Medium / High confidence per locked skill. "Show reasoning" is collapsed by default and expands to the exact source question and quoted answer to provide substantial understanding of the AI reasoning for the skill assessment.

Job requirements list with Met, Partially met, and Not met tags and a line of evidence per requirement
Job requirements

Per-item status, never an aggregate

Each requirement gets its own Met / Partially met / Not met tag with a line of evidence, no "4 of 5" fraction anywhere, so it can't be skimmed as a pass/fail verdict.

05 | Research

Volume outpaced judgment

Customers were asking for faster decisions on high-volume roles, where evaluators had the least time to give each candidate a fair, in-depth read. Alongside my UX design colleague, I analyzed competing products that were already using AI to transcribe, summarize, and evaluate candidate interviews.

We used that research to define our recruiter persona, identify where existing approaches placed AI on the spectrum between information and decision-making, and determine how our feature should be positioned: useful behavioral insights that support a recruiter's judgment rather than replace it.

That tracked with the Challenge: this was built to help evaluators decide, not to be a sales pitch.

High candidate volume

The number of candidates to evaluate kept growing while evaluator time didn't. That tradeoff meant either slow, late evaluations or fast ones that missed what the candidate actually said.

Assisting evaluator judgment

The goal wasn't to replace human judgment, but to support it with objective, evidence-backed insights that removes subjective guesswork and gives evaluators a reason to trust their own call.

Scatter chart plotting reasoning transparency against AI decision authority, with Phenom positioned in the moderate-authority, high-transparency zone alongside HireVue, while Sapia, Workable, BrightHire, and Metaview sit elsewhere

Where this sits against the market

Sapia and Workable score and rank candidates with less traceability behind the number.

HireVue, BrightHire, and Metaview surface evidence but never render a verdict at all.

The shipped direction sits in the narrow band that does both, moderate authority, fully traceable. A spot none of the five actually occupy, since most pick one extreme.

Note: The graph is based on how I read each company's own public product pages, my interpretation of their positioning, not a professional market or product audit.

Working with AI

Problem

Before asking engineering to build this, we needed to know the AI could actually analyze answers against criteria reliably, not just assume it, guess at it, or hope.

AI tools used

Cursor for the prototype, n8n orchestrating the pipeline end to end, Supabase and CloudConvert handling storage and conversion, OpenAI transcribing and analyzing the answers.

Judgement calls

Running it myself surfaced where it breaks: weak prompts, hallucination past a certain number of criteria, vague answers throwing it off. That shaped what we asked engineering to build.

n8n workflow diagram showing the pipeline from screening completion through transcription, OpenAI analysis, and updating the AI Insights record in Supabase
06 | Testing

Four real versions, before landing on the one that shipped

No formal A/B test. Real prototype iterations, each one testing a distinct failure mode rather than a small tweak on the last.

Balanced evaluation version showing candidate summary, advancement/development analysis, core skills confidence, and job requirements
Version A

Balanced evaluation

Shipped

Summary, advancement/development, core skills confidence, requirements

Version with a 1-10 score per criterion, pros and cons, and an overall AI verdict with confidence percentage
Version B

Too much AI authority

1-10 score per criterion, pros/cons, overall AI verdict

Version where the evaluator picks 3 of 10 criteria and generates AI insights on demand
Version C

Evaluator bias

Evaluator picks 3 of 10 criteria, generates on demand delivering inconsistency

Questions and answers view with per-question notes only, plus an AI candidate score with no visible reasoning
Version D

Ungrounded score

Per-question notes only, plus a candidate score with no visible reasoning behind it

Each rejected version fails for exactly one reason, not a vague “we liked A better”:

  • Version B lets a single weak criterion outweigh an otherwise strong candidate, an evaluator can treat one criterion as decisive and reject over one low score.
  • Version C's score means something different depending on which 3 criteria an evaluator happened to pick, so it looks authoritative but isn't comparable across candidates.
  • Version D shows a confident-looking number with nothing on the page explaining where it came from.

Four personas: One weak, one strong, two in between, tested against real AI output across all four versions.

I built a fuller closed-environment prototype with Cursor: a working replica of the product with a questionnaire builder, screening links, and live database. This gave us a realistic environment to run the comparison with real AI output rather than static mockups.

Each teammate submitted their persona's answers through a real screening link, and the AI running on the OpenAI API's latest model at the time, transcribed and analyzed each submission automatically, across all four UI versions.

The comparison wasn't about which layout we preferred. I wanted to know whether the AI's read actually tracked candidate quality consistently, using real output rather than mockups. Building and running the prototype took a few days and gave us evidence before committing engineering effort to the wrong interaction model.

07 | Design decisions

Three real forks

Cohesive vs. per-answer feedbackShipped: Cohesive

Hiring decisions were never tied to one answer. The AI needed to tell one whole-candidate story, not score answers in isolation, which risked bias and didn't map to how the decision actually gets made.

Fit score / candidate comparisonCut from scope

A separate profile-level fit score already existed. A single interview is only one of several stages, scoring off one interview risked losing candidates who'd do better later, and handing evaluators a subjective, biasing number.

Evaluator feedback loopImplicit primary, explicit optional

Confidence per skill is calibrated mainly from behavioral data, decision patterns and downstream candidate progression, learned continuously without requiring anyone to click anything. Explicit thumbs up/down stays optional, with a light prompt only on thumbs-down. Required feedback was the right instinct pointed at the wrong mechanism, coverage matters more than getting everyone to rate.

AI yes/no recommendationCut from scope

Reasons to advance, areas for development, and per-skill confidence already say what's needed. A separate AI verdict next to the evaluator's own Yes/Maybe/No stops assisting and starts voting, least useful on obvious candidates, riskiest on ambiguous ones.

Confidence vs. score - per skillShipped: Confidence

An early version scored each skill 1-10, dropped for putting the AI in the position of judging the candidate. Reframed to measure the AI's own confidence instead, same visual weight, a very different claim.

Job-locked vs. evaluator-picked criteriaShipped: Locked

An earlier version let evaluators pick 3 criteria themselves, per candidate, on demand. Testing surfaced it fast: different evaluators picked different criteria for the same role, and could indirectly steer the AI toward a candidate they already liked. Locking the 3 criteria once, at setup, fixed both problems.

08 | Constraints

A focused product isn't a limited one, it's a positioned one

Cold start

The confidence score had to be built up from evaluator feedback and progression data over time.

Small team

Two designers, one PM, engineering, shaped what got explored vs. shipped inside this phase.

Honest limitation

The real risk isn't a shipped bug, it's evaluators leaning on the score instead of their own judgment, especially early on.

09 | Reflections

Speed is easy, earned trust is not

Building the n8n prototype changed how I think about the role of a designer in AI product development. A working prototype let me test the AI end to end, expose failure modes, and influence product decisions before engineering committed to implementation.

AI can genuinely help evaluators move faster. The open question this feature left me with is whether that speed comes at the cost of blind trust, especially at a stage that feels low-stakes.

Good applied AI looks less like a landing page generated in ten minutes and more like a closed environment built to actually test a feature first.

Last updated on July 23, 2026

Claude Code
LinkedIn
EmailCopy email