hackwither_

Research

Papers

  1. Kaur, B. (2026). Broken Object Level Authorization in the Wild: An Empirical Taxonomy from 100+ Bug Bounty Disclosures. arXiv preprint arXiv:2605.25865.

    Classifies 107 HackerOne disclosures into six BOLA families. Action-level object BOLA (unauthorised state changes on other users’ objects) accounts for 41.7% of confirmed cases.

    107 classified reports: 35 action-level object BOLA, 49 other BOLA families, 23 outside strict criteria84 confirmed BOLA · 78.5%23 out
    Fig. 2 The 107 classified reports, one square each. Pink: action-level object BOLA (35 of 84 confirmed). Solid: the other confirmed families. Outlined: outside strict BOLA criteria.

    200 sampled → 107 classified → 84 confirmed

  2. Kaur, B. (2026). What AI Red-Team Evaluations Can and Cannot Prove. arXiv preprint arXiv:2607.21735.

    Derives a closed-form evidential ceiling: the most a single red-team result can move belief under a fixed testing budget. Benchmarks can certify safety for frequent harms, but fall orders of magnitude short for rare, high-impact ones.

    Schematic: below a calculable harm rate no feasible passive benchmark certifies safety; above it a modest benchmark canrare, catastrophic harmsfrequent harmscrossing · closed formno feasible benchmark certifiesharm rate (log) →
    Fig. 1 Two regimes, schematic and not to scale. An audit of eight evaluation suites finds them adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones.

    discrimination between hypotheses, not attack success, sets evidential worth

Google Scholar profile →

Vulnerability research

  • Hijacking Google ADK Using Malicious A2A Peers

    A remote A2A peer can supply forged control metadata that steers agent routing, session state and even conversation history inside Google’s ADK: a privilege-escalation path from a connected agent to the framework itself.

    Sept 2026 · Google Agent Development Kit
    reported to Google ahead of publication
  • Trust Me, Bro: Forging Security Labels in Microsoft’s FIDES Middleware

    FIDES contains prompt injection by enforcing integrity labels outside the model. Its implementation assumed data can’t describe its own trust level; the finding shows where that assumption breaks and labels can be forged.

    Sept 2026 · Microsoft Agent Framework
    MSRC case 126767 · closed as defense-in-depth
  • Catfishing the Allowlist: MCP Security’s Argument Injection Blind Spot

    An MCP server that allowlists only git still gives full command execution through git’s own alias mechanism. Executable-only allowlists can’t bound what an allowed binary does with its arguments.

    Featured in AI Cyber Magazine, Fall 2026 edition

    Aug 2026 · MCP servers
    GHSA-jm26-853c-62h9 · high
  • Prompt In, Shell Out: Exploiting GitHub AI Toolchains

    Command injection in a GitHub MCP server’s list_issues tool: asking an AI assistant about open issues was enough for remote code execution on the developer’s machine.

    May 2026 · GitHub MCP servers
    coordinated disclosure

Published at APISec Research Labs.

Tools

2026REAP: black-box reconnaissance for AI agent endpoints · 19 rules, MIT, Gogithub
2026reap-range: a deliberately vulnerable MCP target for practice and traininggithub

Analysis

Jun 2026All That Glitters Isn’t Prompt Injection: Rethinking the Meta Instagram Takeover · Meta’s agentic support assistant linked attacker emails to accounts on a one-sentence request. The headlines said prompt injection; the failure was authority.
Jun 2026Beyond the $175K Bankrbot Hack: Encoding Attacks Across Layers of the Agentic AI Stack · The Morse-coded Grok drain was treated as a model-layer problem. Original MCP research shows the same encoding technique working one layer down, at tool execution.
May 2026What 100+ BOLA Reports Taught Us About Where Object-Level Authorization Actually Fails · The blog companion to the BOLA taxonomy paper.

Published at APISec Research Labs.

Selected Disclosures and CVEs

vendorclassstatus
tumf/mcp-shell-serverargument injection → command executionGHSA-jm26-853c-62h9 · high
Model Context Protocol server implementationscommand injection2× CVE reserved
Microsoft Agent Framework (FIDES)security-label forgeryMSRC 126767 · defense-in-depth
Google Agent Development Kitforged A2A control metadatareported
U.S. Department of Education—acknowledged and appreciated by the CISO
Indian government system (reported to NCIIPC)PII and financial data of ~399k government officersacknowledged

I follow coordinated disclosure: vendors are notified privately and given time to fix before details are published. Embargoed work on this page stays redacted until that process is complete.

reserved CVE IDs are linked here once they publish