Ready to See HowThey Really Build?
Recruiter Dashboard / Candidate Comparison
Candidate A
Candidate B
See how they build with AI, not just what they ship.
Request a DemoWhy Now
2015
Engineer writes unaided code
2020
IDE autocomplete and early copilots
2023–2024
Chat-based AI coding assistants
2025–2026
Agentic AI collaboration is the norm
Today
Assessment methodology still reflects 2015
Bellwether
Assessment reflects 2026 engineering
The Thesis
Organizations are hiring 2026 engineers using assessment methodology designed for 2015 engineering.
Practicing engineers in 2026 do not write code unaided; they plan, prompt, review, and refine in continuous dialogue with AI coding assistants. An assessment that forbids the very behavior the job requires is not measuring engineering competence; it is measuring the ability to perform an obsolete version of the job.
The industry is spending its innovation budget defending an assumption that is already false in the field.
Two candidates can submit byte-identical final code: one derived it through careful reasoning and rigorous validation of AI output, while the other pasted the problem into a chatbot and copied the first response. Both receive an identical score today.
Bellwether Whitepaper, §3.4
The Rounds
Six rounds. Three ship in the MVP; three unlock once their question banks are validated.
- 01
Engineering Reasoning
MVPA system-design scenario discussed with an AI collaborator before any code is written — assumptions and constraints articulated out loud.
Design a URL-shortening service for 50 million monthly users.
- Problem Understanding
- Architectural Thinking
- Communication
- 02
AI-Assisted Algorithm Design
MVPThe modern replacement for a traditional unaided DSA round. The AI behaves as an adaptive interviewer rather than an oracle — it asks why this approach, and probes complexity trade-offs.
- Algorithmic Reasoning
- AI Collaboration
- Prompt Engineering
- Testing & Validation
- 03
AI Build Challenge
MVPThe platform's flagship round. An open-ended, realistic build with project scaffolding and package management. All ten dimensions contribute; this round produces the richest evidence trail.
Implement an RBAC-protected inventory API.
- All ten dimensions
- 04
Production Debugging
Production tierAn existing codebase, a bug report, and production-style logs.
- Debugging
- Implementation Quality
- 05
Architecture Review
Production tierIdentify weaknesses in a given architecture across scalability, security, and cost.
- Architectural Thinking
- Problem Understanding
- 06
Engineering Communication
Production tierWritten rationale for earlier decisions: why an architecture was chosen, why an AI suggestion was rejected, what the risk trade-offs were.
- Communication
- Documentation
How It Works
Bellwether turns AI collaboration from a prohibited behavior into the primary signal of engineering competency.
Every prompt, AI response, accepted or rejected suggestion, code revision, and test run is captured as structured signal. The Guardrail Engine keeps the assistant acting as a collaborator rather than an answer key — deterministic rules first, an LLM classifier only as backstop.
The Guardrail Engine's objective is not to prevent AI usage; it is to ensure the AI participates as a collaborator rather than as an oracle.
Guardrail pipeline
Candidate Prompt
Rule-Based Filter
verbatim leakage, direct-answer requests
Redirect — logged as a guardrail event; never a raw error
LLM Classifier
paraphrase & jailbreak detection
Redirect — logged as a guardrail event; never a raw error
Context Injection
round-specific system prompt
LLM Gateway
model-agnostic, multi-provider failover
Response Validation
checked in reverse for leakage
Deliver to Candidate
Gain a Competitive Edge with Bellwether
Builder Score Engine
Ten dimensions, one canonical score. The engine cannot emit a dimension score without an attached evidence object — enforced at the schema layer, so a score missing evidence is a failed transaction, not a warning.
Prompt Intelligence Engine
Classifies intent and scores clarity, context richness, specificity, and constraint definition — then tracks how a candidate's prompts sharpen across a session.
- Technical Execution48%
- AI Collaboration30%
- Professional Judgment22%
- AI Collaboration20%
- Problem Understanding15%
- Architectural Thinking15%
- Prompt Engineering10%
- Algorithmic Reasoning10%
- Implementation Quality10%
- Testing & Validation8%
- Debugging5%
- Communication4%
- Documentation3%
Weights sum to exactly 100% and no dimension may be zero — both enforced as database check constraints, not conventions.
Conversation Knowledge Graph
A candidate's reasoning replayed as a graph of engineering milestones, not a flat chat log: requirement analysis, architecture, API design, implementation, testing, optimization.
Guardrail Engine
Layered and deterministic-first. Rule-based filtering handles the low-latency path; an LLM classifier catches paraphrase and jailbreak attempts only on what clears layer one.
By Design
- 10
- Scored competency dimensions
- 30%
- Of the score legacy platforms cannot see at all
- 100%
- Of scores delivered with an evidence object
- 0
- Autonomous hiring decisions
Exactly ten columns on the score. No more, no fewer.
AI Collaboration and Prompt Engineering — the single largest bloc.
A correctness invariant, not an aspirational target.
A human recruiter always records the decision. Governance, not a feature.
The Difference
Incumbent set: HackerRank, CodeSignal, Codility, LeetCode-style judges
| Traditional OA | Bellwether | |
|---|---|---|
| AI is prohibited | AI is part of the assessment | |
| Only the final code is scored | The reasoning path to the code is scored | |
| A single opaque score is returned | A ten-dimension score, backed by evidence, is returned | |
| Passive dependency on AI is invisible | Passive dependency is explicitly detected and penalized | |
| Cheating is prevented through lockdown | Guardrails shape how AI is used, not whether it is used |
Builder vs. Programmer
The Programmer
Can produce correct code unaided, within familiar problem patterns, under time pressure. This remains a necessary but no longer sufficient competency.
The Builder
Can frame a problem, direct an AI system through planning and delegation, critically validate and refine what it produces, and take full ownership of the resulting system — while retaining the underlying algorithmic and architectural judgment to know when the AI is wrong.
Roadmap
Validate before scale.
The enterprise tier is triggered by a named customer's scale requirements, not by a calendar date. Nothing here is built speculatively ahead of demand.
Phase 1
July 2026
Foundation MVP
PrototypeRounds 1–3, core Builder Score engine, Prototype-tier architecture.
Phase 2
August 2026
AI Depth
PrototypePrompt Intelligence Engine, layered Guardrail Engine, Conversation Knowledge Graph.
Phase 3
September 2026
Pilot Validation
PrototypeStructured pilots with a small number of design-partner organizations. Deliberately the smallest phase by effort, because its purpose is validation, not construction.
Phase 4
October–December 2026
Production Tier
ProductionService decomposition and hardening based on validated pilot learnings.
Enterprise
Customer-triggered
Enterprise Tier
EnterprisePursued only once a specific customer's scale requirements justify it. Not built speculatively ahead of demand.
What We Have Not Proven Yet
No published predictive-validity claim exists yet.
Whether Builder Score components correlate with real on-the-job performance is a hypothesis the Phase 3 pilots and subsequent longitudinal research are designed to test. It is not yet an established fact, and this page does not claim otherwise.
Validate the platform's predictive value with real pilot cohorts before publishing performance or ROI claims, rather than asserting unverified numbers.
Let's Talk
Can this candidate solve a real engineering problem, using the tools a real engineer uses?
Bellwether is pre-pilot and looking for design partners for Phase 3. If you run technical hiring at volume, or certify engineering graduates, that is the conversation.