hackwither_

Bandana
Kaur

aka hackwither

  1. 1Security Research Engineer, APISec Research Labs

Abstract

Hacking a cyberspace that's liveable for all.

19-year-old researcher breaking AI agents: A2A peer hijacking in Google’s Agent Development Kit, security-label forgery in Microsoft’s Agent Framework, reported vulnerabilities in MCP servers with two CVEs assigned. Spoke on three tracks at Black Hat MEA, as one of its youngest speakers: technical briefing, campus keynote and the WiCSME panel. Author of REAP and two arXiv papers.

Portrait of Bandana Kaur

Selected vulnerability research

  • Hijacking Google ADK Using Malicious A2A Peers

    A remote A2A peer can supply forged control metadata that steers agent routing, session state and even conversation history inside Google’s ADK: a privilege-escalation path from a connected agent to the framework itself.

    Sept 2026 · Google Agent Development Kit
    reported to Google ahead of publication
  • Trust Me, Bro: Forging Security Labels in Microsoft’s FIDES Middleware

    FIDES contains prompt injection by enforcing integrity labels outside the model. Its implementation assumed data can’t describe its own trust level; the finding shows where that assumption breaks and labels can be forged.

    Sept 2026 · Microsoft Agent Framework
    MSRC case 126767 · closed as defense-in-depth
  • Catfishing the Allowlist: MCP Security’s Argument Injection Blind Spot

    An MCP server that allowlists only git still gives full command execution through git’s own alias mechanism. Executable-only allowlists can’t bound what an allowed binary does with its arguments.

    Featured in AI Cyber Magazine, Fall 2026 edition

    Aug 2026 · MCP servers
    GHSA-jm26-853c-62h9 · high

Also disclosed: two reserved CVEs in MCP server implementations, records pending publication; a report to the U.S. Department of Education acknowledged by its CISO, and, at 16, a report to India’s NCIIPC covering data of ~399k government officers. All through coordinated disclosure.

record and policy

All findings and analysis →

Keynotes and talks

yearvenuewhere
2026GISEC Global: Exploiting MCP Servers: The AI Let Me In · live hacking demo · also a panel and a fireside chatDubai
2026UNIDIR Global Conference on AI, Security and Ethics (AISE26): Breaking the black box: A standardized lifecycle model for adversarial AI testing · lightning talkGeneva
2026NIST FISSEA Spring Forum: Trust as an Attack Surface: Human Risk in Black-Box AI Systems · speakeronline
2026Gautam Buddha University, Health and Medical Biotechnology Symposium · keynoteGreater Noida
2025Black Hat MEA Campus: The Last Human Hacker: What Comes After AI · keynoteRiyadh
2025Black Hat MEA: Hack one, hack them all? Weaponising LLM jailbreak transferability · technical briefing · also the WiCSME panelRiyadh
All talks and booking →

Selected papers

  1. Kaur, B. (2026). What AI Red-Team Evaluations Can and Cannot Prove. arXiv:2607.21735.

    Derives a closed-form evidential ceiling: the most a single red-team result can move belief under a fixed testing budget. Benchmarks can certify safety for frequent harms, but fall orders of magnitude short for rare, high-impact ones.

    Schematic: below a calculable harm rate no feasible passive benchmark certifies safety; above it a modest benchmark canrare, catastrophic harmsfrequent harmscrossing · closed formno feasible benchmark certifiesharm rate (log) →
    Fig. 1 Two regimes, schematic and not to scale. An audit of eight evaluation suites finds them adequate for high-frequency harm categories and several orders of magnitude short for rare, catastrophic ones.

    discrimination between hypotheses, not attack success, sets evidential worth

  2. Kaur, B. (2026). Broken Object Level Authorization in the Wild: An Empirical Taxonomy from 100+ Bug Bounty Disclosures. arXiv:2605.25865.

    Classifies 107 HackerOne disclosures into six BOLA families. Action-level object BOLA (unauthorised state changes on other users’ objects) accounts for 41.7% of confirmed cases.

    107 classified reports: 35 action-level object BOLA, 49 other BOLA families, 23 outside strict criteria84 confirmed BOLA · 78.5%23 out
    Fig. 2 The 107 classified reports, one square each. Pink: action-level object BOLA (35 of 84 confirmed). Solid: the other confirmed families. Outlined: outside strict BOLA criteria.

    200 sampled → 107 classified → 84 confirmed

All research →

Tool

REAP v0.1.1 · MIT · Go

Black-box reconnaissance for AI agent endpoints. Point it at an MCP endpoint you’re authorised to test: it confirms the protocol, lists what an anonymous caller can reach, and reports auth and transport posture. It reads, never invokes.

Checks per reap-range target: bad 14 matched, 1 clean, 2 skipped; gated 6 matched, 8 clean, 3 skipped; good 1 matched, 13 clean, 2 skipped matched ran clean skipped bad 14/1/2 gated 6/8/3 good 1/13/2
Fig. 3 Every check on three test targets, one square each. Pink: matched. Solid: ran clean. Outlined: skipped, never counted as clean.
brew install hackwither/tap/reap

most scanners tell you what they found. REAP also tells you what it could not check.

Experience

2025–now Security Research Engineer, APISec Research Labs · San Francisco
  • Leads research at APISec Research Labs: original security research on AI systems, agentic workflows, APIs and modern application architectures.
  • Writes the Labs’ public research on real-world vulnerabilities and attack patterns, including the BOLA taxonomy and the agent framework findings.
  • Works with product, engineering and marketing to turn findings into product changes and published research.
2024–now Independent security researcher, hackwither
  • Runs Threat Model Thursdays, a research series on how systems actually get compromised: attack chains, AI misuse, logic flaws, and the failures standard training skips.
  • Coordinated disclosure through public VDPs, including the U.S. Department of Education, and a report to NCIIPC covering data of ~399k Indian government officers.
  • Security education for teenagers on Instagram (70.1K followers) and YouTube; ranked #17 in FeedSpot’s Top 35 Ethical Hacking Influencers of 2026 (up from #26).
2024–now NICE Cybersecurity Career Ambassador (volunteer), NIST
  • Designed and delivered an introductory cybersecurity session for 500+ Gen Z beginners during NICE Cybersecurity Career Week 2024.
  • Ambassador of the Month, April 2025, for live sessions reaching 2,000+ students on cybersecurity careers.
2026 Research Fellow, SPAR (Supervised Program for Alignment Research)
  • Selected for a competitive international alignment research programme (~17% acceptance).
  • Project: proving safety properties of guardrail models.
  • Worked on interpretable constitutional classifiers and formal verification methods for NLP safety systems.
2025 Board Member and Cybersecurity Lead, UN IGF Dynamic Teen Coalition
  • Led the coalition’s cybersecurity channel; appointed to the board in July 2025 to represent teen voices at UN and IGF events.
  • Youth delegate at the IGF 2025 Annual Meeting; took part in the UN Foundation’s Unlock the Future North America brainstorm and UN80 High-Level Week in New York.
  • Panellist on post-quantum cryptography and emerging threats at IS3C, IGF.
  • IYC12 delegate in New York, named Most Engaged Delegate.
2024 Guest lecturer at government CISO training program, Indian Institute of Public Administration · New Delhi
  • Trained 50+ government CISOs, twice, on OSINT and social engineering in the NeGD CISO Training Programme.
  • Awarded IIPA’s special memento on both occasions.
About →

Notes

Jun 2026Acknowledgement is not remediation · At 16 I reported a vulnerability exposing PII and financial data of ~399k Indian government officers to NCIIPC. Finding and acknowledging bugs doesn't secure systems. Remediation does.
Dec 2025The thing about roadmaps in cybersecurity… · How to start in cybersecurity without a one-size-fits-all roadmap: start with the big picture, pick a niche, pivot if it doesn't click, then go deep. No sponsored certs.
All notes →

Work with me

  • The Last Human Hacker: What Comes After AI?

    What comes after AI: how the adversary landscape changes as autonomous agents and synthetic identities enter the threat surface, and why the frontier is human–AI symbiosis. First given as the Black Hat MEA campus keynote.

    keynote · 30–45 min
  • The Hacker’s Guide to AI Agents

    A recon-to-exploitation methodology for agentic systems, MCP and A2A integrations.

    talk · 30 min
  • Evidential Ceilings

    What AI red-team evaluations can and cannot prove. For teams that run, buy or rely on AI safety evals.

    talk · 30–45 min

Keynotes, conference talks or research collaboration: bandana@hackwither.co.in. How to book.