The clock-domain blind spot
Sameer Abu Ghanima · one design, 48 seeds, clock ratios 1:1 to 7:1
We planted a clock-domain-crossing bug in a small handshake design and ran it on 48 seeds. Ordinary RTL simulation passed it every time.3 Then we modeled one thing a real flip-flop can do when it samples a changing signal, settling a cycle late, and the same design failed on 28 of the same 48 seeds.
That gap matters for any grader built on ordinary simulation, including one used to train or evaluate AI models.
The bug
Two parts of the design run on different clocks. A fast side raises requests, and a slow side acknowledges them. The signals between them pass through two-flop synchronizers, the standard way to cross clock domains.
A correct four-phase handshake has the slow side hold its acknowledge until the request drops. In the defective version, the acknowledge is a one-cycle pulse instead. That breaks the return to zero: the fast side can raise its next request before the slow side has seen the previous one go low.
With synchronizers that always take exactly two cycles, the slow side caught that brief low every time across our 48 seeds. But when only one slow-clock edge samples the low, and that flop settles a cycle late, the low is missed: two requests merge into one, and the handshake hangs.
The defective handshake on one real seed: the same clock edges, simulated two ways.
Ordinary simulation: one slow-clock edge samples the brief low between the two requests. The synchronizer captures it on time, it reaches req_sync one cycle later, and the slow side sees two requests.
With late settling: the flop that samples the brief low settles one cycle late. By the next edge req is high again, so the low never reaches req_sync. The two requests merge into one, the second is never acknowledged, and the handshake hangs.
That's how 19 of the 28 failures happen.1 The other 9 all run at clock ratios of 5:3 or closer to 1:1, and there the acknowledge is lost instead. It lasts one slow-clock cycle, so the fast side may sample it on a single edge. When that flop settles late, the acknowledge never arrives. Both mechanisms end the same way: "COMPLETE violated: request N hung".
The measurement
We used the same design and the same 48 seeds (1000 to 1047), covering 13 of the task's 15 clock ratios, from 1:1 to 7:1.4 The run used harness 0.2.0, Icarus Verilog 12 and cocotb 2.0.1.
| Design | Ordinary simulation | With late settling | Caught only with late settling |
|---|---|---|---|
| Defective handshake | 48 of 48 pass | 20 of 48 pass | 28 |
| Correct fix | 48 of 48 pass | 48 of 48 pass | 0 |
| Original clean design | 48 of 48 pass | 48 of 48 pass | 0 |
- No false alarms on this design. The original design and the reference fix, which restores the original logic, pass every seed with late settling on.
- Not one corner. The 28 failures fall at 12 of the 13 clock ratios we ran, from 1:1 to 7:1. The 4 seeds at 7:2 all passed.
- Two simulators. On these 48 seeds, Icarus Verilog 12 and Verilator 5.038 gave the same result and the same trace hash for every seed, in both modes.
The contrast case
We ran a different defect through the same experiment: a logic error in an event-crossing design. It fails all 48 seeds in both modes.
| Design | Ordinary simulation | With late settling |
|---|---|---|
| Event crossing, defective | 0 of 48 pass | 0 of 48 pass |
Late settling doesn't make every bug harder to pass. It exposed the bug that depends on a synchronizer settling late, and it left the logic error's result unchanged.
Why this matters for models that write hardware
A grader built on ordinary simulation can't tell a correct clock-domain fix from one that only works when every flip-flop settles on time. Reinforcement learning optimizes the reward as written, so a timing-blind grader can give full reward to fixes that depend on timing cooperating.
How the late-settling model works
Each synchronizer bit whose input has just changed resolves after two or three destination-clock cycles, chosen independently per bit from the seed. The choices come from the seed, so every result replays exactly. One flag, +CDC_INJECT=0, switches it off, and that's how the comparison above was made. We call this metastability injection. It is a behavioral model, not a circuit simulation.
Limits
- It's a behavioral model. It represents late settling of up to one extra cycle. It doesn't model unknown values, oscillation, longer settling, or failure rates over time.
- Closeness to the edge isn't modeled. The model doesn't depend on how close a change is to the clock edge: any synchronizer bit whose input changed may settle late, at even odds.
- It's one design and 48 seeds. The table is a measurement on that design, not a statement about every design.
- The seeds cover part of the ratio range. The 48 seeds cover clock ratios from 1:1 to 7:1. The task's hidden seeds go up to 7.5:1, and our event-crossing task runs up to 50:1.
Next steps
- If you train models for hardware design or verification, we'll run yours on our clock-domain tasks privately and send you the results, free. Request a free private evaluation
Sameer Abu Ghanima, Sentinely