Bellwether

Ready to See HowThey Really Build?

Recruiter Dashboard

Illustrative

Candidate A

88.6Builder Score

Candidate B

75.4Builder Score
AI Collaboration20%
9560
Problem Understanding15%
9278
Architectural Thinking15%
8895
Prompt Engineering10%
9055
Algorithmic Reasoning10%
8590
Implementation Quality10%
8788
Testing & Validation8%
8070
Debugging5%
8465
Communication4%
8972
Documentation3%
7560
prompt_id · 4a7f2csnapshot_id · 91e0bd

See how they build with AI, not just what they ship.

Request a Demo

Why Now

  1. 2015

    Engineer writes unaided code

  2. 2020

    IDE autocomplete and early copilots

  3. 2023–2024

    Chat-based AI coding assistants

  4. 2025–2026

    Agentic AI collaboration is the norm

  5. Today

    Assessment methodology still reflects 2015

  6. Bellwether

    Assessment reflects 2026 engineering

The Thesis

Organizations are hiring 2026 engineers using assessment methodology designed for 2015 engineering.

Practicing engineers in 2026 do not write code unaided; they plan, prompt, review, and refine in continuous dialogue with AI coding assistants. An assessment that forbids the very behavior the job requires is not measuring engineering competence; it is measuring the ability to perform an obsolete version of the job.

The industry is spending its innovation budget defending an assumption that is already false in the field.

Two candidates can submit byte-identical final code: one derived it through careful reasoning and rigorous validation of AI output, while the other pasted the problem into a chatbot and copied the first response. Both receive an identical score today.

Bellwether Whitepaper, §3.4

The Rounds

Six rounds. Three ship in the MVP; three unlock once their question banks are validated.

  1. 01

    Engineering Reasoning

    MVP

    A system-design scenario discussed with an AI collaborator before any code is written — assumptions and constraints articulated out loud.

    Design a URL-shortening service for 50 million monthly users.

    • Problem Understanding
    • Architectural Thinking
    • Communication
  2. 02

    AI-Assisted Algorithm Design

    MVP

    The modern replacement for a traditional unaided DSA round. The AI behaves as an adaptive interviewer rather than an oracle — it asks why this approach, and probes complexity trade-offs.

    • Algorithmic Reasoning
    • AI Collaboration
    • Prompt Engineering
    • Testing & Validation
  3. 03

    AI Build Challenge

    MVP

    The platform's flagship round. An open-ended, realistic build with project scaffolding and package management. All ten dimensions contribute; this round produces the richest evidence trail.

    Implement an RBAC-protected inventory API.

    • All ten dimensions
  4. 04

    Production Debugging

    Production tier

    An existing codebase, a bug report, and production-style logs.

    • Debugging
    • Implementation Quality
  5. 05

    Architecture Review

    Production tier

    Identify weaknesses in a given architecture across scalability, security, and cost.

    • Architectural Thinking
    • Problem Understanding
  6. 06

    Engineering Communication

    Production tier

    Written rationale for earlier decisions: why an architecture was chosen, why an AI suggestion was rejected, what the risk trade-offs were.

    • Communication
    • Documentation

How It Works

Bellwether turns AI collaboration from a prohibited behavior into the primary signal of engineering competency.

Every prompt, AI response, accepted or rejected suggestion, code revision, and test run is captured as structured signal. The Guardrail Engine keeps the assistant acting as a collaborator rather than an answer key — deterministic rules first, an LLM classifier only as backstop.

The Guardrail Engine's objective is not to prevent AI usage; it is to ensure the AI participates as a collaborator rather than as an oracle.

Guardrail pipeline

  1. Candidate Prompt

  2. Rule-Based Filter

    verbatim leakage, direct-answer requests

    Redirectlogged as a guardrail event; never a raw error

  3. LLM Classifier

    paraphrase & jailbreak detection

    Redirectlogged as a guardrail event; never a raw error

  4. Context Injection

    round-specific system prompt

  5. LLM Gateway

    model-agnostic, multi-provider failover

  6. Response Validation

    checked in reverse for leakage

  7. Deliver to Candidate

Gain a Competitive Edge with Bellwether

Builder Score Engine

Ten dimensions, one canonical score. The engine cannot emit a dimension score without an attached evidence object — enforced at the schema layer, so a score missing evidence is a failed transaction, not a warning.

Prompt Intelligence Engine

Classifies intent and scores clarity, context richness, specificity, and constraint definition — then tracks how a candidate's prompts sharpen across a session.

  • Technical Execution48%
  • AI Collaboration30%
  • Professional Judgment22%
  • AI Collaboration
    20%
  • Problem Understanding
    15%
  • Architectural Thinking
    15%
  • Prompt Engineering
    10%
  • Algorithmic Reasoning
    10%
  • Implementation Quality
    10%
  • Testing & Validation
    8%
  • Debugging
    5%
  • Communication
    4%
  • Documentation
    3%
Total100%

Weights sum to exactly 100% and no dimension may be zero — both enforced as database check constraints, not conventions.

Conversation Knowledge Graph

A candidate's reasoning replayed as a graph of engineering milestones, not a flat chat log: requirement analysis, architecture, API design, implementation, testing, optimization.

Guardrail Engine

Layered and deterministic-first. Rule-based filtering handles the low-latency path; an LLM classifier catches paraphrase and jailbreak attempts only on what clears layer one.

By Design

10
Scored competency dimensions

Exactly ten columns on the score. No more, no fewer.

30%
Of the score legacy platforms cannot see at all

AI Collaboration and Prompt Engineering — the single largest bloc.

100%
Of scores delivered with an evidence object

A correctness invariant, not an aspirational target.

0
Autonomous hiring decisions

A human recruiter always records the decision. Governance, not a feature.

The Difference

Incumbent set: HackerRank, CodeSignal, Codility, LeetCode-style judges

Traditional OABellwether
AI is prohibitedAI is part of the assessment
Only the final code is scoredThe reasoning path to the code is scored
A single opaque score is returnedA ten-dimension score, backed by evidence, is returned
Passive dependency on AI is invisiblePassive dependency is explicitly detected and penalized
Cheating is prevented through lockdownGuardrails shape how AI is used, not whether it is used

Builder vs. Programmer

The Programmer

Can produce correct code unaided, within familiar problem patterns, under time pressure. This remains a necessary but no longer sufficient competency.

The Builder

Can frame a problem, direct an AI system through planning and delegation, critically validate and refine what it produces, and take full ownership of the resulting system — while retaining the underlying algorithmic and architectural judgment to know when the AI is wrong.

Roadmap

Validate before scale.

The enterprise tier is triggered by a named customer's scale requirements, not by a calendar date. Nothing here is built speculatively ahead of demand.

  1. Phase 1

    July 2026

    Foundation MVP

    Prototype

    Rounds 1–3, core Builder Score engine, Prototype-tier architecture.

  2. Phase 2

    August 2026

    AI Depth

    Prototype

    Prompt Intelligence Engine, layered Guardrail Engine, Conversation Knowledge Graph.

  3. Phase 3

    September 2026

    Pilot Validation

    Prototype

    Structured pilots with a small number of design-partner organizations. Deliberately the smallest phase by effort, because its purpose is validation, not construction.

  4. Phase 4

    October–December 2026

    Production Tier

    Production

    Service decomposition and hardening based on validated pilot learnings.

  5. Enterprise

    Customer-triggered

    Enterprise Tier

    Enterprise

    Pursued only once a specific customer's scale requirements justify it. Not built speculatively ahead of demand.

What We Have Not Proven Yet

No published predictive-validity claim exists yet.

Whether Builder Score components correlate with real on-the-job performance is a hypothesis the Phase 3 pilots and subsequent longitudinal research are designed to test. It is not yet an established fact, and this page does not claim otherwise.

Validate the platform's predictive value with real pilot cohorts before publishing performance or ROI claims, rather than asserting unverified numbers.

Let's Talk

Can this candidate solve a real engineering problem, using the tools a real engineer uses?

Bellwether is pre-pilot and looking for design partners for Phase 3. If you run technical hiring at volume, or certify engineering graduates, that is the conversation.