skip to content
// research

September 21st Research Briefing

Contents

  1. ZEBRA: branch-and-bound finds 11 zero-days in production zkVMs
  2. Ghost-filled orders: atomicity gaps in non-custodial prediction markets
  3. Local LLM serving: confidentiality boundaries fail in consumer stacks

Figure 1 from ZEBRA (Takahashi, Jana, and Yang): typical zkVM workflow from source and inputs through instance tables, executor trace tables, prover, and verifier accept or reject.

ZEBRA: branch-and-bound finds 11 zero-days in production zkVMs

ZEBRA verifies zkVM constraint systems by asking whether, for a fixed program and input, the constraints admit exactly one valid execution trace. Too many traces mean an under-constrained soundness bug (forged proofs). Too few mean an over-constrained completeness bug (valid runs rejected). On five production Plonky3 zkVMs (Pico, SP1, Sphinx, Valida, and Ziren) it reports 11 previously unknown zero-days, including control-flow hijacking and proof-forgery flaws. Six were independently confirmed and three were fixed. A state-of-the-art fuzzer found none of them.

This matters because zkVMs move the trust boundary from per-program circuits to a single VM arithmetization. One wrong constraint can admit forged proofs or reject honest execution at production scale.

ZEBRA lifts analysis to an integer interval lattice that exploits measured sparsity (about 14% of theoretical connectivity). Parallel branch-and-bound either emits a concrete counterexample or certifies absence of violations in a bounded region. Against SMT (Z3) it is 51.5× faster on jointly solved instances and verifies 16.5 percentage points more opcode targets (48.0% vs 31.5%).

Authors: Hideaki Takahashi, Suman Jana, and Junfeng Yang (paper, pdf). Hideaki Takahashi: GitHub. Suman Jana: GitHub, LinkedIn.

Ghost-filled orders: atomicity gaps in non-custodial prediction markets

Ghost-Filled Orders studies the gap between off-chain match and on-chain settle in non-custodial prediction markets. An order can be valid when accepted or matched off-chain, then invalid by settlement time after the maker changes nonce, balances, approvals, or receiver hooks. Using Polymarket as the primary case, the paper analyzes 1.8 million reverted transactions touching official contracts over nine months (2025-08-12 to 2026-05-22 UTC), separates attacker profit from market and user loss under conservative rules, and shows adversaries can invalidate unfavorable fills after observing outcomes or price moves.

This matters because hybrid CLOB designs keep user custody while still exposing a classic TOCTOU window. Batched settlement means one ghost fill can revert a whole bundle and hurt unrelated traders.

The authors then apply the same audit methodology to three additional blockchain prediction markets. All three were exploitable; one design also enables direct attacker profit. Findings were reported to those projects, and one had acknowledged the issue at study time.

Authors: Zhiyang Chen, Fan Long, and Zhendong Su (paper, pdf). Zhiyang Chen: GitHub, X, LinkedIn. Fan Long: X. Zhendong Su: GitHub, LinkedIn.

Local LLM serving: confidentiality boundaries fail in consumer stacks

The Illusion of Local Privacy shows that running an LLM on a local machine does not by itself keep prompts confidential. Prompt lifetime depends on model load, runtime memory, wrapper persistence, and the serving interface. LLAnalyzer tests each boundary on four open-weight families (Nemotron-3-Nano, Qwen3.5, Gemma-4, and Phi-4-reasoning-plus) and two consumer platforms (LM Studio and Ollama).

This matters because “local inference” is often treated as the end of the privacy story. In practice confidentiality is a composition of software layers, and a single weak boundary can undo the rest.

Runtime memory retains recoverable plaintext after inference, and wrappers can extend lifetime through plaintext persistence. At the serving boundary the authors uncover a previously undocumented authorization flaw in llama.cpp: an authenticated client can restore another tenant’s saved conversation state in 200/200 controlled trials. Separately, shared prompt-prefix caching yields a remote timing oracle that remains distinguishable under WAN conditions. Artifact: Local-LLMs-Forensics.

Authors: Youssef Hamdi Zafan Ibrahim, Muhammad Ikram, and Mohammed Khalaf Salama (paper, pdf, Local-LLMs-Forensics). Muhammad Ikram: GitHub, LinkedIn.

// related
September 28, 2026// research
September 28th Research Briefing

A weekly publication of the top three papers from arXiv last week. Papers with arXiv v1 published 2026-09-21 through 2026-09-28 UTC.

September 14, 2026// research
September 14th Research Briefing

A weekly publication of the top three papers from arXiv last week. Papers with arXiv v1 published 2026-09-07 through 2026-09-13 UTC.

Research notes, at most monthly. No spam.
← All research