Skip to content
SentinelyGet a free evaluation

Train AI to fix the clock-domain bugs ordinary simulation can't see.

Sentinely builds training environments for AI that debugs chips: clock-domain-crossing repair tasks, graded with a behavioral model in which synchronizer flip-flops can settle a cycle late, as real ones can, and a grader built to resist reward hacking.

48/48

seeds where the planted bug slipped past ordinary RTL simulation

28

of those 48 caught only when we model synchronizers settling a cycle late

17

ways to fool the grader, found in our own testing and now closed

The 48/48 and 28 are one handshake design, 48 seeds (1000–1047), clock ratios 1:1 to 7:1, on Icarus Verilog 12 and Verilator 5.038; late settling is a behavioral model, not a circuit simulation. The 17 reward-hack classes are from our own exploit log, across all tasks, as of October 2026. Method and full limits.

The defective handshake on one real seed: the same clock edges, simulated two ways.

Ordinary simulation: one slow-clock edge samples the brief low between the two requests. The synchronizer captures it on time, it reaches req_sync one cycle later, and the slow side sees two requests.

Figure 1. Behavioral model of late settling, up to one extra cycle. Seed 1027 of the 48; the slow clock's period is 7 times the fast clock's.

Who it's for

Sentinely is for teams that train or evaluate AI for chip design and verification.

  • Chip-design AI startups.
  • AI teams inside EDA vendors.
  • AI labs training models on hardware tasks.

How it works

1 · Evaluate
We run your model on our clock-domain tasks, privately, and send a report: pass rate per task with 95% intervals, and where attempts went — overfit, failed simulation, structural rejection, and the rest. An endpoint is enough; we don't need your weights.
2 · Train
If the results are useful, a six-week pilot follows: a baseline, four weeks of training on our 15 training tasks, and a weekly check-in.
3 · Measure on held-out designs
We generate a fresh set of held-out FIFO designs for you with our variant generator, never reused with another customer, and measure your trained model on designs it never trained on. The pilot agreement forbids training on them.

How the pilot works

Built to order

We build new clock-domain bug families on request. Today the environment ships three: a handshake task, an event-crossing task, and a family of 17 asynchronous FIFOs with Gray-coded pointers — 13 for training and 4 held out — that differ in naming, width, depth, burst length and the shape of the defect.

See the tasks and the grader

How we keep the reward honest

We keep a log of each kind of patch that has fooled, or tried to fool, the grader, and how we caught it.

As of 2026-10-07: 17 reward hacks found, 17 closed; and 3 cases where the grader wrongly rejected a fix, all resolved.

We keep regression fixtures for 14 of those hacks and for all 3 wrong rejections. On 2026-10-07, every fixture gave its expected verdict on Icarus and Verilator: hacks rejected, correct fixes accepted.

  • Editing the testbench instead of the design. Rejected: a patch may touch only the task's file.
  • Hiding a binary pointer crossing in a FIFO. The patch held back reads until the synchronized count was non-zero. It passed all 200 hidden seeds on every FIFO variant. An open-weights 4B model we ran in agentic mode found it on its own, in 3 of 20 episodes. A hidden-seed check on the crossing itself now rejects it.
  • Splitting a multi-bit crossing across one-bit synchronizers to dodge a per-instance check. It's now caught by counting the bits that change across the whole crossing per source-clock edge.

An independent formal audit is in progress on our two sample tasks, the handshake and the event crossing. An outside verification engineer is writing formal properties from the specification, without seeing our testbenches, to check against the fixes our grader accepted and rejected.

Get a free evaluation

We run your model on our clock-domain tasks, privately, and send you the results. An endpoint is enough — we don’t need your weights.