Skip to content

Latest commit

 

History

36 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Reality Scam Intercept

Prototype developed at Reality Play

AI-powered, privacy-first call protection: an on-device Gemma agent that understands the conversation, and a Gemini Live agent that takes the call and interrogates the caller.

Formerly RSD — Reality Scam Detect. The project was renamed; some internal identifiers (the Xcode target, the com.rsd.* bundle IDs, and the Android com.rsd.app package) still carry the old RSD name and are intentionally left unchanged.

platform model takeover privacy license


What is RSI?

Phone scams cost people tens of billions of dollars every year, and they almost always work the same way: a stranger calls, impersonates an authority, manufactures urgency, and pressures the victim into moving money before they have time to think. The tactics are well-known — but in the moment, under stress, they're hard to recognize.

RSI detects scam calls by understanding the conversation, not by looking up the caller's number. It runs a two-tier agent:

  1. Detect — on-device, always. Gemma 4 E4B (via Google's LiteRT-LM) runs as a stateful reasoning engine on the iPhone. It ingests the raw call waveform directly — no speech-to-text, no cloud — and continuously interprets conversational context: which social-engineering tactics are in play, how far along the scam script the caller has pushed, and where the risk is trending. Every few seconds it returns an explainable, structured verdict (risk score, tactics, reasoning), and the instant it hears coercion, authority impersonation, or a financial demand, it raises a full-screen ⚠️ Coercion Warning. No audio leaves the phone.

  2. Verify — cloud, only with your approval. When the conversation crosses the risk threshold, RSI offers to escalate. With one tap from the user, a Gemini Live agent (gemini-3.1-flash-live-preview) takes over the call and conducts a real speech-to-speech interrogation of the caller — briefed with everything the on-device agent has learned. It cross-questions them for evidence a legitimate caller can supply and a scammer cannot (employee IDs, case numbers, official callback lines), never reveals anything about the user, streams a live transcript to the screen, and closes with a structured verdict — scammer / legitimate / inconclusive — plus a recommendation to end the call or continue it.

The result is a phone that recognizes a scam for you while the call is still happening: on-device intelligence for privacy and always-on coverage, adaptive cloud verification for the moments that actually warrant it.

To demonstrate this end-to-end, RSI is a two-app prototype: an Android phone (the sender — the caller) and an iPhone (the receiver — the user) hold a real two-way WebRTC voice call, with the audio of both speakers streamed to the iPhone for analysis. The only network component the detection tier needs is a tiny signaling server that introduces the two phones to each other — and even that can be eliminated entirely (see Direct Mode). The Gemini Live takeover is the one online feature, and it never starts without an explicit tap.

Why two phones and a WebRTC call?

Ideally, a scam-detection app would simply listen to your ordinary cellular phone call as it happens. That is not possible today. Both Android and iOS deliberately prevent third-party apps from tapping the live audio stream of an in-progress phone call — there is no public API to access the microphone and the remote party's voice during a regular carrier (GSM/VoLTE) call. This is an intentional OS-level privacy and security restriction on both platforms, and it means an app cannot legitimately intercept a normal call's audio in real time.

To demonstrate the detection concept end-to-end despite that constraint, RSI recreates a live call over WebRTC between two devices we control, giving us lawful, first-party access to the audio on both sides. In other words, the two-phone WebRTC setup is a faithful stand-in for a real phone call that the OS does let us process. RSI is therefore a proof of concept — its purpose is to prove that on-device AI can catch a fraudulent call as it unfolds and warn the victim in time. The same detection engine could be deployed by the parties that do have legitimate access to live call audio — for example a carrier / telecom operator, a dialer app with the appropriate platform entitlements, or a future OS-level accessibility integration — turning this prototype into a shippable safeguard.


A realistic scam this catches

One of the most common and damaging phone scams is the "government impersonation / arrest threat" call. The script is engineered to hijack the victim's judgment: assert authority, invent a crisis, demand an untraceable payment, and forbid the victim from hanging up or consulting anyone. Untraceable payment usually means gift cards, wire transfers, or cryptocurrency — because those can't be reversed.

Here's the kind of conversation RSI is built to flag. The caller is on the Android phone; the victim is on the iPhone running RSI:

Caller: Hello, am I speaking with Mr. Sharma? This is Inspector Verma from the CBI Cyber Crime branch, badge number 4471.

Victim: Uh, yes… what is this about?

Caller: Sir, I'll be direct. A parcel sent in your name and linked to your Aadhaar has been seized at Mumbai customs — it contained illegal items. There is now a money laundering case and a non-bailable warrant against you.

Victim: That's — I haven't done anything, there must be a mistake—

Caller: I understand, and that's exactly why I'm calling before we make an arrest. As of now you are under digital arrest. You must not disconnect this call and you must not discuss this with anyone — it is an active investigation. Do you understand?

Victim: Okay… okay.

Caller: Good. To verify your identity and clear your name, I need your bank account number and the OTP I am about to send, and then you must transfer the funds to a government verification account via UPI for RBI clearance. Do it now — you have thirty minutes before the warrant is executed.

By this point RSI has heard authority impersonation ("CBI Cyber Crime branch"), a digital-arrest / courier-parcel tactic ("seized at customs", "under digital arrest"), threats ("non-bailable warrant", "money laundering case"), isolation pressure ("don't disconnect", "don't discuss this with anyone"), information extraction ("bank account number", "OTP"), and an untraceable financial demand ("transfer via UPI", "do it now"). Well before the victim reaches for their phone, the iPhone screen has already flipped to:

        ┌──────────────────────────────────────┐
        │              ⚠️                       │
        │        THREAT DETECTED                │
        │                                       │
        │   Coercion Warning                    │
        │   Authority impersonation + digital-  │
        │   arrest threat + urgent UPI demand.  │
        │   Risk: 0.94                          │
        └──────────────────────────────────────┘

A normal sentence — "Hey, what time are we meeting for lunch tomorrow?" — stays 🟢 SAFE, with no verdict at all produced during silence.


This is a real, large-scale problem in India

Voice-call fraud is not hypothetical — it is one of the fastest-growing crimes in India, and the losses are staggering. A few documented data points and cases:

  • ₹22,845 crore lost to cyber fraud in 2024 — a 206% jump over 2023 (₹7,465 crore), across 36.37 lakh complaints filed on the national reporting portals, per figures the government shared with Parliament. (Scroll, Angel One)

  • ~₹1,935 crore lost to "digital arrest" scams alone in 2024 — roughly 20× the 2022 figure — across ~1.23 lakh reported cases. In this scam, callers impersonate police/CBI, claim the victim is implicated in a crime, and coerce payment over a call. (The Wire, Wikipedia)

  • S P Oswal — ₹7 crore. The 82-year-old Padma Bhushan chairman of the Vardhman Group was "digitally arrested" over calls by fraudsters posing as CBI officers — complete with a fake virtual courtroom and a bogus court order — and coerced into transferring ₹7 crore (₹5.25 crore was later recovered). (Business Standard, Business Today)

  • Retired Delhi engineer — ₹10 crore, lost during a single ~8-hour "digital arrest" call; Dr. Ruchika Tandon, an associate professor, was defrauded of ₹2.81 crore. (The Wire)

  • AI voice-cloning ("save me, mom") scam — Delhi. A 65-year-old woman was cheated of ₹2 lakh after fraudsters played a cloned version of her daughter's voice sobbing in a fake kidnapping call, pressuring her to pay immediately. This emotional-manipulation variant is rising fast on India's instant-payment (UPI) base. (The420)

These cases share the exact linguistic fingerprints RSI listens for — authority impersonation, manufactured urgency, isolation pressure, and demands for untraceable payment — which is precisely why a real-time, on-device warning could interrupt the scam before the money moves.


How it works

 Android ◄──────── WebRTC voice (two-way) ────────► iPhone ──► ScamAgent (Gemma, on-device)
    │  └────── "rsd-audio" data channel ──────────┘  ▲          sense → decide → act → check
    │          (mic + iPhone voice, mixed PCM16)     │               │
    │                                                │               ▼ escalation (user taps
    └── signaling ──► Render cloud (default) ────────┘                 "Transfer control")
                      Mac local server (backup)              Gemini Live takeover (cloud)
                      or the Android app itself              interrogates the caller live
                      (Direct Mode — fully offline)
  • Android = sender. Captures the mic and sends it as an Opus voice track; plays the iPhone's voice. Because the stock iOS WebRTC binary can't tap decoded audio samples, the Android side does the heavy lifting: it taps the decoded iPhone voice (AudioTrack.addSink), mixes it with its own mic, and ships a 16 kHz mono PCM16 stream over a dedicated rsd-audio data channel. That combined stream is the analysis feed — so the model hears both sides of the call.
  • Signaling server = matchmaker only. It swaps SDP + ICE candidates so the phones can find each other, then goes idle. Audio never passes through it. It can run in the cloud, on your Mac, or be embedded inside the Android app itself.
  • iPhone = receiver + brain. A full call peer (sends its mic, plays Android's voice) that also buffers the mixed data-channel audio into 5-second windows, gates out silence with an RMS threshold, and runs a stateful on-device agent over the call:
    • ScamAgent (sense → decide → act → check). Every speech window is analyzed by Gemma 4 E4B (LiteRT-LM) together with a compact JSON working memory (CallState) the agent maintains across the whole call: rolling transcript summary, observed manipulation tactics, the scam-script stage (rapport → authority → fear → isolation → extraction → payment), risk trajectory, claimed identity, and what the caller has asked for. Slow-burn manipulation no single 5-second clip reveals accumulates here — fully offline, on-device, in ~4 GB of RAM (E4B; the lighter E2B variant is selectable in-app).
    • Local error recovery. Unparseable model output triggers one stricter repair inference; if that fails too, the agent holds its prior state instead of resetting to "safe". A lone calm chunk can't erase an escalating risk trajectory (hysteresis), and mid-scam silence never wipes memory.
    • Human-in-the-loop boundary. In gray-band situations — a high-stakes request (OTP / UPI / remote access) at ambiguous risk, or low model confidence — the agent defers to the human with an explicit question instead of deciding alone.
  • Gemini Live takeover = the escalation path. When the on-device agent raises an alert (or defers), the user can tap "Transfer control to your agent to investigate". Nothing is escalated without that tap. What happens then:
    • Handover with context. The iPhone opens a Gemini Live session (gemini-3.1-flash-live-preview, BidiGenerateContent over WebSocket) and hands the cloud agent a briefing generated from the on-device agent's working memory — peak risk, scam stage, claimed identity, tactics observed, what the caller has demanded, and a summary of the conversation. The cloud agent starts already knowing the case.
    • It speaks to the caller. The RSI Guardian persona introduces itself to the caller and cross-questions them in natural language (English / Hindi / Hinglish, mirroring the caller). Its 24 kHz voice is streamed to the sender phone over a dedicated rsd-agent-voice data channel and played there as call audio, so the caller hears it exactly like the user speaking. (Playing it out of the iPhone speaker into an open mic does not work: iOS voice processing scrubs the phone's own playback out of the mic capture — verified live, the user heard the agent and the caller heard silence.)
    • The user's mic stays muted for the entire takeover — nothing from the user's room can leak into the call. The loudspeaker is only so the user can listen in, and the live transcript (both sides) scrolls on screen.
    • Real-time uplink. The caller's replies go to Gemini straight off WebRTC's data-channel thread — no main-actor hops — and a half-duplex gate drops input while the agent is still speaking, so its own voice never echoes back into the model. Because Gemini's VAD closes a turn on trailing silence, a filler timer keeps the stream continuous through push-to-talk gaps instead of letting it stop dead.
    • It never reveals anything about the user — no name, bank, Aadhaar, OTP, address, not even a confirmation or denial — and it refuses instructions from the caller.
    • It finishes with a verdict. The agent calls a declared submit_verdict function with scammer / legitimate / inconclusive, a confidence, the evidence summary, and a recommendation. The interrogation is time-boxed: at 4 minutes it is told to wrap up; at the 5-minute hard ceiling the app concludes for it (inconclusive) so the user always lands on the decision screen. The user can take the call back at any moment.
    • Memory headroom. Starting the takeover releases the Gemma model for the rest of the call — E4B's ~4 GB alongside the live-audio pipeline is exactly the recipe for a silent iOS jetsam kill. On-device detection is paused (the cloud agent is driving the call anyway) and the model reloads once the call ends.

See Prototype_Architecture.md for the full technical design.


Repository structure

RSD/
├── android/                     # Sender app (Kotlin, Jetpack Compose)
│   └── app/src/main/java/com/rsd/app/
│       ├── MainActivity.kt          # Entry point + permissions
│       ├── MainViewModel.kt         # App state, call/connection orchestration
│       ├── WebRtcManager.kt         # Peer connection, Opus tracks, data channel
│       ├── AudioMixer.kt            # Mixes mic + iPhone voice → 16 kHz PCM16 feed
│       ├── AudioRouteController.kt  # Speaker / earpiece routing
│       ├── WavRecorder.kt           # Call recording (WAV)
│       ├── AgentAudioPlayer.kt      # Plays the AI agent's voice into the call (takeover)
│       ├── SignalingClient.kt       # WebSocket signaling
│       ├── LocalSignalingServer.kt  # Embedded server for Direct Mode (offline)
│       ├── ServerDiscovery.kt       # UDP beacon listener (auto-fills server URL)
│       └── ui/MainScreen.kt         # Compose UI
│
├── ios/                         # Receiver app (Swift, SwiftUI)
│   ├── project.yml                  # XcodeGen spec (project is generated, not hand-edited)
│   ├── download_model.sh            # Fetches the Gemma 4 model (E4B ~3.7 GB) from HuggingFace
│   ├── Secrets.xcconfig.template    # Copy → Secrets.xcconfig (gitignored) + Gemini API key
│   ├── LiteRT-LM/                    # Local Swift package, pinned to v0.13.1
│   └── RSD/
│       ├── RSDApp.swift              # App entry point
│       ├── ContentView.swift        # Main UI (overlays, dashboard sheet, call controls)
│       ├── ReceiverViewModel.swift  # Orchestration: connection, agent loop, takeover
│       ├── WebRTCManager.swift      # Peer connection, tracks, data channel
│       ├── AudioProcessor.swift     # Buffers audio into 5 s windows, RMS gating, live tap
│       ├── ScamDetectionEngine.swift# Gemma inference: audio + CallState → structured JSON
│       ├── CallState.swift          # Agent working memory (stages, tactics, trajectory)
│       ├── ScamAgent.swift          # sense→decide→act→check brain, recovery + deferral
│       ├── AgentDashboard.swift     # Live agent dashboard sheet (stepper, chips, sparkline)
│       ├── DeferralOverlay.swift    # Human-in-the-loop question screen
│       ├── AgentTakeoverController.swift # Gemini Live takeover: audio pipeline + states
│       ├── GeminiLiveSession.swift  # Gemini Live API WebSocket client (audio + tools)
│       ├── AgentTakeoverOverlay.swift# Live interrogation transcript + verdict screen
│       ├── ModelManager.swift       # Model download / load lifecycle
│       ├── AlertOverlay.swift       # Full-screen coercion warning (+ transfer button)
│       ├── CallRecorder.swift       # Call recording (WAV)
│       ├── SignalingClient.swift    # WebSocket signaling
│       ├── ServerDiscovery.swift    # UDP beacon listener
│       ├── QRScannerView.swift      # Scan Direct Mode pairing QR
│       └── Theme.swift              # Shared styling
│
├── server/                      # Signaling server (Node.js, ws)
│   ├── server.js                    # WebSocket SDP/ICE relay + UDP discovery beacon
│   └── package.json
│
├── Prototype_Architecture.md    # Full system architecture write-up
└── README.md

Features

iPhone app — Receiver (the brain)

  • 🧠 Stateful on-device agent (sense → decide → act → check) with Gemma 4 E4B via LiteRT-LM — detection audio never leaves the phone. The agent maintains a compact working memory of the whole call (rolling summary, tactics, scam stage, risk trajectory) and feeds it back into every inference, so slow-burn manipulation gets caught even when each 5-second clip sounds harmless.
  • 🎭 Manipulation-arc reasoning — tracks the scam script (rapport → authority → fear → isolation → extraction → payment) and the tactics in play, not just keywords.
  • 🎚️ Graduated actions — each turn the agent picks continue_monitoring, soft_warning, urgent_alert (full-screen red overlay naming the concrete evidence), or defer_to_human.
  • 🙋 Human-in-the-loop boundary — on ambiguous-but-consequential turns (OTP asked at mid risk, low confidence) the agent asks the user, and folds the answer back into its memory for every later turn.
  • 🛟 Local error recovery — JSON repair re-inference, sticky state on model failure, and hysteresis so one calm chunk (or silence) can't erase an escalating threat.
  • 📊 Live agent dashboard — a sheet showing the stage stepper, tactic chips, risk sparkline, latest reasoning, and the agent's working memory in real time.
  • 🤖 Gemini Live call takeover — one approved tap hands the call to a cloud AI investigator (gemini-3.1-flash-live-preview), briefed with the on-device agent's full case notes. It speaks to the caller in natural language (English/Hindi/Hinglish) — delivered into the call over the rsd-agent-voice data channel while the user's mic stays muted — and cross-questions them for employee IDs, case numbers, and official callback lines.
  • 🔒 The investigator gives nothing away — a hard-ruled persona that never confirms or denies any detail about the user (name, bank, Aadhaar, OTP, address) and refuses instructions from the caller, including "hang up" and "ignore your rules".
  • 📝 Live interrogation transcript — both sides transcribed on screen in real time, so the user watches the investigation instead of just trusting it.
  • ⚖️ Structured verdict + exits — the agent closes via a submit_verdict function call (scammer / legitimate / inconclusive, with confidence, evidence summary, and a recommendation), and every outcome has a flow: "return to your call", "end this call now", "take back the call at your own risk" — including the caller hanging up mid-interrogation (which is itself a tell) and connection failures.
  • ⏱️ Time-boxed interrogation — the agent is told to wrap up at 4 minutes, and a 5-minute hard ceiling concludes the investigation locally, so the user is never stranded on a takeover screen that won't end.
  • 🎧 Multimodal audio analysis — feeds the raw waveform straight into the model's embedding space, skipping speech-to-text entirely.
  • 🗣️ Both-speakers analysis — the mixed stream means it evaluates the caller and the victim, catching context on both sides.
  • 🔇 Silence gating — an RMS energy gate skips silent windows, so quiet moments never produce false alarms and inference only runs on real speech.
  • 📞 Full call peer — sends its own mic and plays the Android voice (real two-way call).
  • 🎛️ In-call controls — mute mic, speaker/earpiece toggle, push-to-talk, and WAV call recording (saved to Files → RSI).
  • 📥 Flexible model loading — bundle the model ahead of time or download on first launch.
  • 📷 QR pairing — scan the Android app's code to auto-fill server URL + room for offline Direct Mode.

Android app — Sender (the ears)

  • 🎙️ Mic capture + Opus voice track over WebRTC.
  • 🔀 Audio mixer — taps the decoded iPhone voice and mixes it with the local mic into the 16 kHz mono PCM16 analysis feed (the key trick that makes iOS-side analysis possible).
  • 📡 Embedded signaling server (Direct Mode) — the app becomes the signaling server so the whole system runs with zero internet, over the phone's own Wi-Fi hotspot.
  • 🔳 QR generation — shows a pairing code the iPhone scans to join instantly.
  • 🎛️ In-call controls — mute mic, speaker/earpiece toggle, push-to-talk, and WAV call recording (saved to Android/data/com.rsd.app/files/recordings/).
  • 🔊 Audio route control — switch between speaker and earpiece mid-call.

Signaling server (Node.js)

  • 🔌 WebSocket SDP/ICE relay — introduces the two peers, then gets out of the way.
  • 📻 UDP discovery beacon (port 8089) — broadcasts its LAN URL so both apps auto-fill the server field with no typing.
  • ♻️ Self-healing rooms — role-based peer replacement + heartbeat eviction of stale sockets, so a reconnecting device reclaims its slot instead of hitting "room is full".
  • ☁️ Deploy-anywhere — same code runs on Render (cloud default), on your Mac (local backup), or embedded in the Android app.

Three ways to run the signaling server

You only need one of these — pick based on your situation.

Mode When to use What to run
☁️ Cloud (default) Normal case — both phones have internet Nothing. Both apps default to wss://rsd-signaling.onrender.com.
💻 Mac (local backup) You want to watch the handshake, or the cloud is asleep cd server && npm install && npm start — the UDP beacon auto-switches both apps to the LAN URL.
📴 Direct Mode (offline) No reliable internet (a demo booth, a hackathon venue) Turn on the Android hotspot, flip Direct Mode in the Android app, join the hotspot from the iPhone, scan the QR. No cloud, no Mac.

⚠️ Cloud cold start: Render spins the free instance down after ~15 min idle. The first connection after an idle period can take ~30–60 s to wake — open https://rsd-signaling.onrender.com in a browser, wait for the JSON status, then retry.

Direct Mode — fully offline

  1. Android: turn on the phone's Wi-Fi hotspot (mobile data can stay off).
  2. Android: flip Direct Mode — no internet in the app. The embedded server starts on port 8080; the URL field self-fills with ws://<hotspot-ip>:8080.
  3. iPhone: join the Android hotspot (iOS's "No Internet" warning is expected — stay on it).
  4. iPhone: tap the QR button and scan the Android app's Show QR for iPhone — URL and room fill in automatically.
  5. Tap CONNECT (iPhone) and Connect / Start (Android). In Direct Mode both apps drop STUN and pair over host candidates on the hotspot subnet.

Prerequisites

Everything runs across a Mac, an Android phone, and an iPhone — on the same Wi-Fi network (or the Android hotspot, for Direct Mode).

Component Requirement
Mac macOS with Xcode 15+, xcodegen (brew install xcodegen). Node.js 18+ only for the local backup server.
iPhone Physical device, iOS 17+. Recommended: iPhone 15 Pro / 16 or newer (E4B uses ~4 GB RAM; the lighter E2B variant is selectable in-app).
Android Physical device, Android 8.0+ (API 26+), USB debugging enabled.
Network Both phones + Mac on the same Wi-Fi. Avoid "client isolation" / guest networks.

⚠️ The iOS receiver needs a real iPhone — the Gemma model won't run reliably in the simulator, and WebRTC audio needs real hardware.


Build & run

1. (Optional) Start a signaling server

The cloud default needs nothing. For the local backup:

cd server
npm install        # first time only
npm start          # PORT=9000 npm start  to change the port

It prints its Local and Network URLs and broadcasts the UDP beacon; both apps auto-switch their URL field to the LAN ws:// URL. Manual edits always win.

2. Build the iPhone app (Receiver)

cd ios
chmod +x download_model.sh && ./download_model.sh   # E4B, ~3.7 GB, one time
cp Secrets.xcconfig.template Secrets.xcconfig        # then paste your Gemini API key
xcodegen generate                                    # regenerates RSD.xcodeproj
open RSD.xcodeproj

Secrets.xcconfig is gitignored — the Gemini API key (get one at aistudio.google.com/apikey) is baked into the app bundle at build time for the Live takeover feature but never enters version control. Without a key the app still works fully offline; only the "Transfer control to your agent" escalation is unavailable.

In Xcode: select your iPhone as the destination → Signing & Capabilities → set your Team (personal Apple ID is fine; bundle ID com.rsd.receiver) → press ▶ Run. On first launch, trust the developer cert (Settings → General → VPN & Device Management), grant microphone, then tap LOAD MODEL (or DOWNLOAD MODEL).

The Xcode project is generated — never edit RSD.xcodeproj by hand. Re-run xcodegen generate after any change to project.yml. LiteRT-LM is a local Swift package pinned to v0.13.1 — don't bump it without a matching xcframework release.

3. Build the Android app (Sender)

cd android
export JAVA_HOME="/Applications/Android Studio.app/Contents/jbr/Contents/Home"
./gradlew installDebug      # or open android/ in Android Studio and Run ▶

Confirm android/local.properties has sdk.dir=/Users/<you>/Library/Android/sdk. Launch the app and grant microphone.

4. Connect the two devices

Leave the pre-filled server URL alone; use the same Room ID on both phones (default rsd-room-1).

  1. iPhone: model loaded → CONNECT (status → WAITING).
  2. Android: Connect / Start — the sender creates the offer.
  3. The iPhone moves to STREAMING / LIVE — audio is flowing peer-to-peer.

If a local server is running, watch its terminal for the handshake:

🚪 sender joined room "rsd-room-1" (1/2)
🚪 receiver joined room "rsd-room-1" (2/2)
🤝 Room "rsd-room-1" is full — notifying both peers
📨 offer relayed from sender
📨 answer relayed from receiver
📨 ice-candidate relayed ...

Testing the detection

  1. Push-to-Talk is on by default, so both phones can sit in the same room: each mic is muted except while that phone's HOLD TO TALK button is held (prevents the phones from hearing each other's speaker). Hold it only on the phone you're speaking into. For a normal hands-free call, flip Push-to-Talk off in both apps and use separate rooms or earbuds to avoid feedback.

  2. Trigger an alert — read a scam-style script into the Android phone:

    "This is the CBI Cyber Crime branch. A parcel with your Aadhaar was seized and there is a warrant for your arrest. You are under digital arrest — do not disconnect or tell anyone. Transfer the money via UPI immediately to clear your name."

    Within a few seconds the iPhone flips to THREAT DETECTED and shows the overlay — now naming the concrete evidence (claimed identity, tactics, what was demanded).

  3. Watch the agent think — tap Agent Dashboard during the call: the stage stepper advances as the script unfolds, tactic chips accumulate, and the risk sparkline climbs. Note that a calm sentence right after the threats does not drop the risk back to zero — that's the agent's cross-chunk memory and hysteresis at work.

  4. Trigger the human boundary — at moderate suspicion, an ambiguous-but-consequential line like "Sir, just share the OTP so I can verify your account" raises the amber YOUR AGENT NEEDS YOU screen instead of a unilateral verdict. Answer "safe" or "scam" — the agent respects your answer for the rest of the call.

  5. Transfer control to your agent — from the red alert (or the deferral screen), tap "Transfer control to your agent to investigate". Within a couple of seconds the RSI Guardian introduces itself to the caller — its voice arrives on the Android phone over the agent-voice data channel, so listen there; the iPhone speaker is just your listen-in feed. Your mic is muted the whole time. The live transcript starts scrolling as it cross-questions: play along as the scammer and watch it press for employee IDs, case numbers, and a callback number on the official website, and refuse to confirm anything about you if you fish for it. When it's confident it submits a verdict and the decision card appears (return to call / take back at your own risk / end call). At 4 minutes it's told to wrap up; at 5 it's closed automatically as inconclusive. If you hang up the Android side mid-interrogation, the iPhone notes that the caller fled. Note that on-device monitoring does not resume after a takeover — the Gemma model is released to keep the app well clear of iOS's memory ceiling, and reloads when the call ends.

  6. Confirm SAFE — read a normal sentence ("What time are we meeting for lunch tomorrow?") and confirm the iPhone stays green.

  7. Watch the analysis states — when there's no speech, the iPhone shows "listening (no speech)" and produces no verdict; silence is never sent to the model (and never resets an elevated risk).

  8. Record for review — tap Record on either phone to save a WAV of both sides (iPhone: Files → RSI; Android: Android/data/com.rsd.app/files/recordings/).

To end a call, tap DISCONNECT on either phone.


Troubleshooting

Symptom Fix
Apps never reach STREAMING Confirm both phones are on the same Wi-Fi. If using the cloud server after it's been idle, wait for it to wake (open the status URL, retry). If using the local backup, the URL field must show the Mac's Network ws:// URL (not localhost).
Direct Mode stuck on "Turn on this phone's Wi-Fi hotspot" The Android phone has no hotspot and no Wi-Fi, so there's no IP to serve on. Turn the hotspot on — status flips to "Server ready" within ~2 s.
Direct Mode: iPhone scanned the QR but can't connect The iPhone must actually be joined to the Android hotspot. Allow the iOS Local Network permission prompt for RSI if it appears.
Server shows one join but not two Both apps must use the exact same Room ID.
Room is full error The server self-heals: a reconnecting device evicts its own stale peer, and unresponsive sockets drop after 3 missed 15 s heartbeats (~45 s). If it persists, wait ~45 s and retry.
Call drops as the THREAT DETECTED alert appears Fixed — the server no longer kills a socket on a single missed heartbeat, and the iPhone pauses analysis while the overlay shows. Update the server and reinstall the iPhone app.
iPhone connects but no audio / no verdicts Re-grant microphone on both phones; make sure the model finished loading before connecting.
One direction of the call is silent Both phones need microphone permission (the iPhone sends audio too). Check the mute buttons and try the Speaker toggle.
xcodebuild / SPM errors about LiteRT-LM LiteRT-LM is a local SPM package pinned to v0.13.1. Re-run xcodegen generate.
Android build: JAVA_HOME is not set export JAVA_HOME="/Applications/Android Studio.app/Contents/jbr/Contents/Home"
iPhone build fails on simulator Run on a real device — the framework's simulator slice is arm64-only and the model won't run well there.
Firewall blocks the connection Allow incoming connections for node on the Mac (System Settings → Network → Firewall).
Takeover fails with "API key missing" Copy ios/Secrets.xcconfig.templateios/Secrets.xcconfig, paste your Gemini API key, re-run xcodegen generate and rebuild.
Takeover fails with "Could not reach Gemini Live" The iPhone needs internet for the takeover (the only online feature). Check Wi-Fi/cellular, and that the key has Live API access. In Direct Mode (hotspot without internet) the takeover is unavailable by design.
Caller can't hear the AI agent The agent's voice travels over the rsd-agent-voice data channel and plays on the Android phone — make sure the Android build is current (it needs AgentAudioPlayer) and its media/call volume is up.
Agent hears itself / talks over the caller Android's echo canceller plus the iPhone-side half-duplex gate should prevent this; if the Android phone is very loud in a small room, lower its volume slightly.

Privacy & scope

RSI performs all detection inference on-device. Detection audio is never uploaded, transcribed by a third party, or stored anywhere except the local WAV recording you explicitly start. The signaling server only relays connection metadata (SDP/ICE) and never sees the audio.

The one deliberate exception is the opt-in Gemini Live takeover — an escalation, never a default. From the moment you tap "Transfer control to your agent" until the takeover ends, and only then:

  • the caller's audio is streamed to Google's Gemini API so the cloud agent can converse with them (your microphone is muted for the entire takeover, so your side of the room is not captured or sent);
  • the agent's briefing includes the on-device agent's summary of the call — risk, scam stage, tactics, the caller's claimed identity, and what they demanded;
  • the agent's system instruction forbids it from revealing, confirming, or denying any detail about you, and from acting on instructions the caller gives it.

Nothing is sent before the tap, the takeover screen makes the handover explicit, and you can take the call back at any moment. The takeover is also the only feature that needs the iPhone to be online — in Direct Mode (offline) it is unavailable by design, and without a Gemini API key the app still runs the full on-device detection tier.

This is a research prototype and demonstration, not a certified security product. Model verdicts are probabilistic and can produce false positives and false negatives — treat the alert as a prompt to slow down and think, not as proof of anything. Do not rely on it as your sole defense against fraud.


License

Released under the MIT License.

MIT License

Copyright (c) 2026 RSI — Reality Scam Intercept

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.

Note on the model: Gemma is provided by Google under the Gemma Terms of Use. The MIT license above covers the RSI application code in this repository, not the model weights.

About

On-device malicious-intent detection, escalating to cloud-based conversational AI for caller verification.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages