Reviewer lens: is this worth building, for one technical owner, on his own Samsung (maybe Pixel), and what is the smallest thing that proves it?
Inputs: PRD-original.md (truncated at §7) and DESIGN-DRAFT.md. Date: 2026-10-03.
Bottom line. The PRD is written like a consumer product heading for the Play Store. It carries Play-policy constraints, consumer-grade transparency settings and five numeric dimensions. The owner is one person sideloading onto his own phone. Most of the scope protects against problems he does not have.
The draft's claimed differentiator is the stateful trajectory (opener → wrong-number pivot → platform shift → investment). As of 2026 that is exactly what Google's on-device Scam Detection targets, and it is on by default for non-contacts. The genuinely unique value is narrower:
That is a 2-week build, not a multi-phase program. Gate it on a measured volume of unknown-sender messages before writing the classifier.
For the narrow core value ("don't call a stranger spam just because they're a stranger"), mostly yes. For the PRD's full ambition (history-aware, multi-signal trust), no.
The estimates below are reviewer judgment, not measurements. The Phase 0 probe should replace them with real numbers.
| Missing data | What it costs | Est. share of total value lost |
|---|---|---|
| Outbound history ("I texted this number first") | This is the single strongest legitimacy signal for a non-contact: the plumber you texted, the clinic you replied to last month. Without it, every unsaved number you have already engaged with looks like a cold stranger. The MessagingStyle tail (android.messages) sometimes includes your recent replies in the current thread, which is partial and unreliable. | 20–30%. The biggest loss. |
| Spam-folder messages | Google posts no notification for messages it files as spam, so the sidecar never sees them. This is the PRD's core thesis failing the other way: the worst "unknown treated as malicious" error (a legitimate stranger silently auto-filed) is invisible to the sidecar. The sidecar can only re-grade what Google already let through. | 10–20%, and it is the most painful 10–20%. |
| Full thread history | Weakens trajectory detection. It is mitigated if the sidecar keeps its own per-sender verdict history from the notifications it has seen, which the draft already proposes. | 5–10% |
| MMS images / attachments | Image-only scams (QR codes, screenshot "invoices") get "Cannot assess". These are rare in personal SMS volume. | <5% |
| Verified-business / RCS metadata | Loses a strong positive signal for business senders. Short codes partially substitute. | ~5% |
Net: notification-only keeps maybe 50–65% of the conceivable value. Almost all the loss is on the legitimacy side (history), not the risk side. That is backwards for this product, which exists to protect legitimacy.
Cheap recovery the PRD forbids by inheritance, not necessity. READ_SMS and READ_CALL_LOG are banned by Google Play policy, not by the OS. For an adb-installed personal app they may be grantable. Mark this verify in the spike, because hard-restricted permission allowlisting and Android 13+/15 "restricted settings" may still gate it. Two caveats:
READ_SMS yields SMS/MMS outbound history only.READ_CALL_LOG ("I called or was called by this number") is a strong, cheap legitimacy signal and covers the RCS gap partially.This contradicts PRD §4.2. It is labeled as an owner decision: "the non-goal exists because of Play policy, which does not apply to a personal sideload; revisit?"
Cheap partial fix for spam-folder blindness (on-thesis, high value): if an expected sender (dentist, FedEx, new employer) is due and no matching message has arrived by the expected time, nudge: "Expected Bright Smile Dental — nothing arrived. Check Spam & blocked in Messages." That is the only way a notification-only sidecar can rescue a stranger Google filed away.
| Layer | What it does today | Source |
|---|---|---|
| Google Messages Scam Detection | On-device AI flags conversational scams (job, romance-baiting / pig-butchering, "switch to another app") across SMS/MMS/RCS. On by default, only for non-contacts, English, US/UK/CA. The user can dismiss, or report and block. No explanation of why is described. | https://blog.google/security/new-ai-powered-scam-detection-features/ |
| Gemini Nano upgrade to the above | Flagships (Pixel 10 series, Galaxy S26, select others): "more nuanced analysis" of gradual-manipulation conversations. | https://9to5google.com/2026/02/25/google-messages-scam-detection-gemini/ |
| Google Messages link / sender protections | Warnings on links from unknown senders, an unknown-international-sender filter (to Spam & blocked), and Key Verifier for contact identity. These were announced 2024 and rolled out in stages by region. | https://security.googleblog.com/2024/10/5-new-protections-on-google-messages.html and https://blog.google/technology/safety-security/how-google-protects-against-scams-2025/ |
| Classic Google spam filter | Auto-files spam to "Spam & blocked" with no notification, which is the blind spot in §1. | same |
| Samsung Message Guard | A zero-click image sandbox, not scam-text triage. Irrelevant to this problem. (Samsung's "AI spam filter" headlines refer to a Korea KCC program and do not apply to US Galaxy.) On a US Samsung the real competitor is Google Messages. | https://docs.samsungknox.com/admin/fundamentals/whitepaper/samsung-knox-mobile-security/application-security/samsung-message-guard/ |
| Carrier filters | T-Mobile network-side SMS/MMS/RCS spam fingerprinting plus Scam Shield. Verizon/AT&T similar. Reporting via 7726. These are blunt and opaque, and they kill bulk spam before it reaches the phone. | https://www.t-mobile.com/scamshield |
| Android OS | Android 15 redacts OTP-like notifications from untrusted listeners (RECEIVE_SENSITIVE_NOTIFICATIONS is role/signature-gated). This is good for privacy and bad for analysis completeness. | https://www.androidauthority.com/android-15-sensitive-notifications-3416414 |
What this means for the draft's "differentiator":
Residual problem only this app solves:
Everything else in the PRD duplicates work done upstream.
Ranked by value ÷ effort:
| # | Feature | Value | Effort | Verdict |
|---|---|---|---|---|
| 1 | Phase 0 probe as a volume meter (contact / non-contact / short code per week, from Person uri and title format only, no text) | Kill gate | 1 day | Keep. Do first. |
| 2 | Listener + MessagingStyle parse + "in contacts" inferred from Person uri (no READ_CONTACTS) | Foundation | 2 days | Keep |
| 3 | ~15 deterministic reason codes + the 6-verdict decision table + golden corpus tests | Core | 3 days | Keep, with verdict + codes only |
| 4 | Expected-sender list (free text + optional brand/number + expiry) | Highest unique value | 1–2 days | Keep |
| 5 | Annotation notification (silent, verdict + top 2 reasons) | Core UX | 1 day | Keep |
| 6 | "Ask who this is" smsto: draft | High, tiny | 0.5 day | Keep |
| 7 | "Expected sender didn't arrive → check Spam" nudge | High, on-thesis | 0.5 day | Keep |
| 8 | Trust sender (local list) | Medium | 0.5 day | Keep |
| 9 | Last-50 log screen (verdict + codes, no bodies) | Medium (debug + trust) | 1 day | Keep |
| 10 | Per-sender last-5 verdict history (trajectory-lite) | Low–medium (overlaps Google) | 1 day | v1.1 |
| 11 | Five numeric dimensions shown in the UI | Low. Verdict + codes carries the meaning, and completeness collapses into "Cannot assess" | 2 days | Cut (compute internally if needed) |
| 12 | Quiet-unknowns re-notify mode | High if right, and dangerous | 3–4 days | Cut from v1. Its failure mode (sidecar dead, so messages from real people go silent) is the exact harm the product exists to prevent. Revisit after 30 days of annotate-mode precision data. |
| 13 | Feedback-learned weight multipliers | Low at n≈single-digit/week | 2 days | Cut. Manual trust plus editing the rules file is enough for one technical user |
| 14 | On-device LLM (Gemini Nano / ML Kit) | Low, redundant with Google's own Nano | 4+ days | Cut |
| 15 | Calendar matching (READ_CALENDAR) | Medium | 1–2 days | v1.1 |
| 16 | SQLCipher + HMAC key rotation | Low for a no-body, no-INTERNET personal app | 1–2 days | Cut. App-private storage, no bodies, no INTERNET permission is already strong |
| 17 | Work-profile / managed-device handling, multi-app ingestion | ~0 for this owner | — | Cut |
| 18 | Deep-link via stored contentIntent + BAL opt-in | Medium | 1 day | Simplify: smsto: fallback only |
Total kept ≈ 10–11 working days. That fits 2 weeks with slack for Samsung battery-optimization fights.
READ_SMS + READ_CALL_LOG on a sideloaded build (§1). It recovers most of the outbound-history value without abandoning Google Messages. Verify grantability in the spike.Keep it to six numbers, read from the local log, reviewed at day 30 (go/no-go) and day 60:
A soft signal: does Byron open the annotation at least weekly by week 4? If not, the tool has become furniture. Kill it, or fold the expected-sender idea into something he already looks at.