Reviewer lens: does the two-ledger log-odds + decision-table design produce verdicts that are right, explainable, and stable? Are the reason codes grounded in how US smishing actually looks in 2024–2026?
Companion artifacts (all in this folder):
golden-corpus.json: 42 cases. Each has expected_verdict, must_include_codes, must_not_include_codes, expected_codes_full, and reference_scores.tools/score.py: a reference implementation of the scoring rules in §1.5 (priors, family caps, decision table). It is the spec the Kotlin core must match.tools/detect.py: a reference implementation of every detector in §2, in stdlib Python. Its regexes are canonical. The ones in the markdown tables below escape | as \| and are for reading only.tools/corpus.py runs the detectors on the raw message text and thread context, then scores what they emit. It also diffs the emitted codes against expected_codes_full, checks must_not_include_codes, and checks the verdict. Run it with python3 tools/corpus.py. Current result: 0 verdict mismatches, 0 code diffs, 42 cases. The first detector pass found real false positives, including a bank alert's merchant descriptor "BESTBUY.COM" parsed as a link and tripping BRAND_DOMAIN_MISMATCH. Those were fixed in the rules and are written up in §2.The architecture is right: two ledgers, reason codes with published weights, and a decision table instead of a single threshold. Keep it. It has five defects that will cause real misclassifications, and each one is fixable in a line or two of spec:
Fix: add negation guards, require imperative-plus-object, and restrict the brand table to high-impersonation brands.
Everything below makes these concrete.
The ledgers split "who is this" from "what is this asking". That split is correct and it is the product's core differentiator. The missing axis is who controls the evidence:
| Attacker-controlled (cheap to fake) | Not attacker-controlled (hard evidence) |
|---|---|
| Self-identification text ("this is X from Y") | Sender is in Contacts (Person URI is a contacts lookup) |
| Brand names in text | User marked the sender trusted |
| "New number" claims | User previously sent a substantive reply in this thread |
| Relationship claims ("we met at…") | Short code in the bundled brand table |
| Using the user's first name (it is in data-broker dumps) | Self-ID matches a specific entry in the private context ledger |
| Calendar match | |
| Every link resolves to the named brand's allowlisted eTLD+1 | |
| Sender has ≥3 prior clean messages over ≥7 days |
Rule: soft codes may move a message within Unrecognized. Only hard codes can lift it to Likely legitimate. In the reference implementation this is the hard flag. Corpus pair G22 vs G23 tests it: identical Carlos text, with and without a ledger entry.
Two subtleties:
Naive-Bayes summing assumes the codes are conditionally independent. They are not. "late fee", "$6.99", "pay by 10/04" and "toll" all come from the same scam template. Fix:
| Family | Members | Aggregation |
|---|---|---|
| LINK | URL_*, LOOKALIKE_DOMAIN, BRAND_DOMAIN_MISMATCH | sum, cap +3.0 |
| ASK | OTP / gift card / crypto / remote-access / credential / payment / callback | sum, cap +3.0 |
| PRETEXT | toll, package, bank, job, govt, carrier, crypto-invest, prize | max (one pretext per message) |
| PRESSURE | urgency, penalty, small fee | sum, cap +1.0 |
| SENDER | brand-from-long-code, intl-claims-US-entity, group blast | sum, cap +1.0 |
| SOCIAL | opener, wrong-number pivot, platform shift, new number, family emergency, secrecy | sum, cap +2.0 |
| EVASION | reply-to-enable-link, injection, homoglyph | sum, cap +2.5 |
The caps keep a single template from saturating the score through repetition. Independent families still add. A toll smish with a lookalike link and the reply-Y trick should be at about 100, and is (G16, L = +8.5).
The prior carries the base rate by sender kind:
| sender_kind | prior logit | scam_risk with no codes | Rationale |
|---|---|---|---|
| contact / user-trusted | −4.0 | 2 | Compromise is possible but rare. A separate gate handles it (§1.5). |
| short code (5–6 digits) | −2.5 | 8 | Carrier-vetted (CTIA short code program). Scam short codes exist but are uncommon. |
| toll-free (8XX) | −2.0 | 12 | Verification is required for toll-free A2P, but enforcement is weaker. |
| US long code | −1.5 | 18 | Mixes personal numbers and 10DLC business numbers. |
| international | −0.8 | 31 | Smishing Triad uses foreign SIMs. Real friends abroad also text. |
| email-to-SMS gateway | −0.5 | 38 | Heavily abused by toll/USPS kits. A few legitimate uses (legacy alerts, people texting from email). |
| sender unavailable | −1.5 | 18 | Also lowers completeness. |
All values are starting points (§5). G38 (a +44 friend) and G39 (a bare "Hello, are you free?" from an email gateway) pin the requirement that no sender kind alone reaches Suspicious (40). G39 sits at 38 on purpose: if calibration pushes the email-gateway prior above about −0.4, that test fails loudly, and that is the intended alarm.
Important: the sidecar sees a post-filter distribution. Google Messages' own spam filter suppresses notifications for what it catches, so the sidecar sees what Google missed plus all ham. The priors therefore should be lower than "fraction of unknown texts that are scams", and public datasets (full of easy spam) overstate how separable the residual is.
As drafted, legitimacy = identity + context − risk is redundant. It is a mix of three other displayed numbers, and once risk is folded in it becomes a noisy second copy of 100 − scam_risk. That is the exact conflation the PRD forbids ("unknown is not malicious").
The PRD requires five dimensions, so do not drop it. Redefine it so it carries information the others do not:
100·σ(−2 + Σ identity codes).100·σ(−2 + Σ context codes).100·σ(−2 + Σ identity + Σ context), capped at 59 when no hard code is present. It never includes risk. It is the "positive case for this message" number, and it is displayed but is not an input to the decision table. The table reads identity, context and risk directly.100·σ(prior(sender_kind) + Σ family-capped risk codes + negative codes). Negative codes are allowed only for hard evidence: KNOWN_SHORT_CODE_BRAND (−1.0) and LINK_DOMAIN_VERIFIED (−1.0). Context codes never lower risk. This is the draft's own rule, and it is the right one.The point of keeping legitimacy and risk as separate quantities: both can be high at once, and that is the most informative state the product can show. Two examples:
A single score cannot say either sentence.
UI consequence for contacts. For contacts, scam_risk with a −4 prior stays low (G24 reports 23) while the verdict says Suspicious. That will confuse users. Either display a content risk computed with the neutral long-code prior (−1.5) next to sender trust, or show only the triggered codes for contacts. Recommended: always compute both scam_risk (sender-aware, used for gating) and content_risk (prior −1.5, for display and for the compromise gate).
The draft checks completeness first, which hides visible tripwires. Use this ordering instead:
G33 ("Your account has been tempor…") must land on Unrecognized, not Likely legitimate.
tools/score.py)no visible text (completeness < 40) -> Cannot assess
sender is contact / user-trusted:
tripwire OR content risk-code sum >= 2.5 -> Suspicious + POSSIBLE_ACCOUNT_COMPROMISE
(Likely scam/phishing if scam_risk >= 70)
else -> Recognized
any tripwire -> Likely scam/phishing if scam_risk >= 70, else Suspicious
scam_risk >= 70 -> Likely scam/phishing
scam_risk >= 40 -> Suspicious
(identity >= 60 OR context >= 60) AND >=1 hard code AND risk < 40
-> Likely legitimate
otherwise -> Unrecognized / insufficient context
Logit thresholds: 40 ↔ L = −0.405 and 70 ↔ L = +0.847.
Differences from the draft:
Text normalization runs first and is shared by every detector:
hxxp→http, [.]/(.)/ dot →., and usps . com → usps.com.URL extraction:
\bhttps?://[^\s<>"]+\b(?:[a-z0-9](?:[a-z0-9-]{0,61}[a-z0-9])?\.)+[a-z]{2,24}\b(?:/[^\s]*)?Columns: weight is in log-odds. T = tripwire. H = hard evidence (attacker cannot control).
| Code | Rule | Prior | FP notes | ||||||
|---|---|---|---|---|---|---|---|---|---|
| UNKNOWN_SENDER | No contacts-lookup URI in android.people.list and not in the local trust list | weight 0 (prior carries it) | Never decisive, by construction. | ||||||
| (kind) short code | Sender is ^\d{5,6}$ | −2.5 | Some 3–4 digit carrier codes (e.g. 728, 611): treat ^\d{3,6}$ as short code. | ||||||
| (kind) toll-free | `^\+?1?(800\ | 833\ | 844\ | 855\ | 866\ | 877\ | 888)\d{7}$` | −2.0 | |
| (kind) long code | ^\+?1?[2-9]\d{2}[2-9]\d{6}$ | −1.5 | 10DLC business numbers live here too. | ||||||
| (kind) international | E.164 with country code ≠ +1, or NANP non-US (+1 876/809/… Caribbean) | −0.8 | Caribbean +1 area codes are a classic one-ring/"wangiri" vector. They look domestic, so keep a list. | ||||||
| (kind) email gateway | Sender contains @, or is a known gateway pseudo-number | −0.5 | Some real alerts (old school/county systems) send via email-to-SMS. Hence a prior, not a tripwire. |
| Code | Wt | H | Rule | FP / abuse notes | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CONTACT_MATCH | +5.0 | H | Person.uri starts with content://com.android.contacts/contacts/lookup/, or the android.title display name is non-numeric with a lookup URI. Verify in the spike. | Spoofed caller-ID SMS is rare but possible. Hence the compromise gate. | |||||||||
| USER_TRUSTED | +5.0 | H | HMAC(normalized number) is in the local trust list | Behaves like a contact. | |||||||||
| THREAD_USER_REPLIED | +2.5 | H | android.messages tail contains a message with sender == null/self whose text has ≥3 words or ≥15 characters and is not in the trivial-reply set | Pig-butchering victims also reply substantively. That is why this is identity, not a risk reducer. | |||||||||
| THREAD_USER_REPLIED_TRIVIAL | 0 | The user's only reply matches `^\s*(y\ | yes\ | n\ | no\ | 1\ | ok\ | k\ | stop\ | wrong number\ | who is this\??)\s*$` | Exists to block reply-Y laundering. | |
| SHORT_CODE | +1.0 | H | Sender kind is short code | ||||||||||
| KNOWN_SHORT_CODE_BRAND | +2.0 id, −1.0 risk | H | Short code is in the bundled {code → brand} table and the brand named in the text (if any) equals that brand | The table must come from brand-published sources. Short codes are country-specific. | |||||||||
| SENDER_HISTORY_CLEAN | +1.0 | H | Per-sender state: ≥3 prior messages, first seen ≥7 days ago, no risk code ≥1.0 ever | Slow-burn pig-butchering builds clean history. Cap at +1.0 for that reason. | |||||||||
| SELF_IDENTIFIED | +0.7 | `\b(this is\ | it'?s\ | i'?m\ | my name is)\s+([A-Z][a-z]+(\s[A-Z][a-z]+)?)(\s+(from\ | with\ | at\ | of)\s+([A-Z][\w&'.-]+(\s[A-Z][\w&'.-]+){0,4}))?, or a leading ^[A-Z][\w &'.-]{2,40}:\s` brand prefix | Scammers self-identify constantly (G19, G26). Soft by design. | ||||
| LINK_DOMAIN_VERIFIED | +1.5 id, −1.0 risk | H | All URLs' eTLD+1 are in the allowlist of the brand named in the text, or match a domain the user attached to a ledger entry | A real-brand link can be paired with a callback-number vishing ask. CALLBACK_NUMBER still fires. |
| Code | Wt | H | Rule | FP notes | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EXPECTED_SENDER_SPECIFIC | +2.5 | H | The self-ID entity, or a ≥2-token proper-noun span, fuzzy-matches (Jaro-Winkler ≥0.92) a named non-expired ledger entry ("Bright Smile Dental", "Carlos" + "Ashford") | Common first names alone ("Carlos") must not match. Require entity + org, or a full name. | |||||||||||||||
| EXPECTED_SENDER_GENERIC | +1.0 | Text names a carrier/brand category that matches a ledger entry ("FedEx", "UPS", "Amazon") | Soft: scammers send carrier lures by the million. Never lowers risk (G09). | ||||||||||||||||
| CALENDAR_MATCH | +1.5 | H | Org name or attendee in text matches an event within ±3 days (optional READ_CALENDAR) | Calendar invites from scammers exist. Require the org to match the event title, not just the description. | |||||||||||||||
| THREAD_CONTINUITY | +1.5 | MessagingStyle tail has ≥1 prior exchange in both directions, and the new message is topically continuous (shared nouns/time references) | Same reply-Y caveat. Requires a non-trivial user reply. | ||||||||||||||||
| APPOINTMENT_REMINDER | +0.8 | `\b(appt\ | appointment\ | visit\ | reservation\ | booking)\b and a date/time \b(mon\ | tue\ | wed\ | thu\ | fri\ | sat\ | sun\ | \d{1,2}/\d{1,2})\b.*\b\d{1,2}(:\d{2})?\s?(am\ | pm)\b and reply\s+(c\ | y\ | 1\ | confirm)` | Scam "appointment" lures exist (fake DMV appointments). Soft weight. | |
| OTP_DELIVERY | +0.5 | \b\d{4,8}\b near `(code\ | otp\ | passcode\ | pin) and a negation (do not\ | don'?t\ | never)\s+(share\ | give) or (will never ask)` | Drop the digit span before storage. Usually redacted on Android 15+ anyway. | ||||||||||
| IN_PERSON_DELIVERY | +0.5 | `(outside\ | at (the\ | your) (door\ | gate\ | lobby)\ | downstairs\ | which (building\ | unit\ | door))` and package/delivery wording, and no URL | |||||||||
| ADDRESSES_USER_BY_NAME | +0.3 | The user's configured first name appears in the text | Data-broker leaks make this cheap. Keep it tiny. |
| Code | Wt | T | Rule | FP notes |
|---|---|---|---|---|
| URL_PRESENT | +0.3 | ≥1 URL extracted | Real reminders contain links. Tiny on purpose. Merchant descriptors are not links: a scheme-less domain right after at/from in a purchase/charge/transaction context ("purchase at BESTBUY.COM") is skipped. Without that rule, every real bank alert for an online purchase trips BRAND_DOMAIN_MISMATCH (found by running the corpus, G06). | |
| URL_SHORTENER | +0.7 | eTLD+1 in {bit.ly, tinyurl.com, t.co, is.gd, rb.gy, cutt.ly, ow.ly, shorturl.at, tiny.cc, s.id, t.ly, …} | Brand-owned shorteners (e.g. amzn.to, vz.to): put them in the brand allowlist, not here. | |
| URL_RISKY_TLD | +1.0 | TLD in {top, xyz, icu, cyou, vip, shop, live, info, click, buzz, sbs, cfd, bond, online, lat, win, rest, mom, lol, help, support, today} | .shop/.live/.info have real users. That is why it is not a tripwire. Refresh this list from SmishTank/Unit 42 TLD stats. | |
| URL_IP_LITERAL | +1.5 | Host matches ^\d{1,3}(\.\d{1,3}){3}$ or contains [ (IPv6) | Almost no legitimate SMS uses IP literals. | |
| URL_OBFUSCATED | +0.8 | De-fanging changed something, or the URL is split across spaces/line breaks to defeat linkification | ||
| LOOKALIKE_DOMAIN | +2.5 | T | Any of: (a) xn-- punycode label; (b) a brand/agency token from the brand table appears in a label of the host but the eTLD+1 is not on that brand's allowlist (ezpass.com-tollpay.top, usps.redelivery-care.cyou, wellsfargo.secure-verify-login.com); (c) eTLD+1 second-level label within Damerau-Levenshtein ≤2 of a brand label (fedx.com, uspss.com) after confusable mapping | (b) needs a brand table with care: "chase" in chasebank-realtors.com? Restrict tokens to ≥4 characters and to high-impersonation brands. Expect some FPs on fan/affiliate sites; the tripwire only forces Suspicious, not Likely scam. |
| BRAND_DOMAIN_MISMATCH | +2.5 | T | Text names a brand from the high-impersonation table (USPS, UPS, FedEx, DHL, Amazon, Apple, Google, Microsoft, PayPal, Venmo, Zelle, Cash App, Chase, BofA, Wells Fargo, Citi, Capital One, Amex, Navy Federal, USAA, IRS, SSA, DMV, E-ZPass, SunPass, FasTrak, TxTag, Peach Pass, Toll By Plate, Netflix, T-Mobile, Verizon, AT&T, Coinbase, …), and some URL's eTLD+1 is not on that brand's allowlist | Excluded on purpose: small businesses, clinics and dentists, which routinely use vendor domains (Weave, Solutionreach, NexHealth, Klara, Podium; G04). Marketing click-trackers (e.g. click.e.brand.com) usually share the eTLD+1; if not, add them to the allowlist. |
| Code | Wt | T | Rule | FP notes | |||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OTP_SHARE_REQUEST | +3.0 | T | `(send\ | give\ | tell\ | forward\ | text\ | read)\s+(me\ | us\ | it\ | back)?.{0,20}\b(the\ | that\ | your\ | a)?\s*(\d[- ]?digit\s+)?(code\ | otp\ | pin\ | passcode\ | verification) or (what('?s\ | is) the code\ | accidentally sent .{0,30}code). Negation guard: suppress if within 40 characters of (do not\ | don'?t\ | never\ | will never ask\ | no one from)` | Real 2FA texts say "never share" (G11, must_not_include). |
| GIFT_CARD_PAYMENT | +3.0 | T | Imperative buy/pay verb `(buy\ | purchase\ | pick up\ | get me\ | pay (with\ | in\ | using)) within 40 characters of (gift\s?cards?\ | itunes\ | apple card\ | google play\ | steam\ | ebay card\ | vanilla\ | razer gold), or (send\ | scratch).{0,30}(codes?\ | pins?\ | numbers on the back)` together with gift-card words | Retail promos ("get a $10 gift card when you buy 2") must not match. Require the card to be the payment instrument (G27). | |||||
| CRYPTO_PAYMENT_DEMAND | +3.0 | T | `(send\ | deposit\ | transfer\ | pay\ | top up) within 40 characters of (btc\ | bitcoin\ | usdt\ | tether\ | eth\ | crypto\ | wallet address\ | bitcoin atm), or a BTC/ETH/TRON address regex \b(bc1\ | [13])[a-km-zA-HJ-NP-Z1-9]{25,39}\b\ | \b0x[a-fA-F0-9]{40}\b\ | \bT[1-9A-HJ-NP-Za-km-z]{33}\b` | Users who trade crypto get real exchange notices. Those come from known short codes and don't demand deposits. | |||||||
| REMOTE_ACCESS_APP | +3.0 | T | `(install\ | download\ | open) within 40 characters of (anydesk\ | teamviewer\ | quicksupport\ | rustdesk\ | screenconnect\ | ultraviewer\ | zoho assist)` | Legitimate IT support by SMS is vanishingly rare for consumers. | |||||||||||||
| CREDENTIAL_REQUEST | +1.5 | `(verify\ | confirm\ | update\ | validate\ | restore)\s+(your\s+)?(identity\ | account\ | card\ | details\ | information\ | login\ | ssn\ | social security)` | Real banks say "log in to the app to review". Hence not a tripwire. | |||||||||||
| PAYMENT_REQUEST | +1.0 | `\b(pay\ | payment\ | settle\ | balance due\ | outstanding (amount\ | balance)\ | fee)\b` with an imperative or a deadline | Real bills (G31) carry it. Neutralized by LINK_DOMAIN_VERIFIED + KNOWN_SHORT_CODE_BRAND. | ||||||||||||||||
| CALLBACK_NUMBER | +0.8 | A phone number in the text plus `(call\ | contact)` in an alert/fraud context | Real fraud alerts say "call the number on the back of your card", which has no number. Explicit numbers in alerts are a vishing tell. | |||||||||||||||||||||
| REPLY_YES_NO_ALERT | +0.3 | `reply\s+(yes\ | y)\s+or\s+(no\ | n)` in a transaction-alert context | Real banks do exactly this (G06). Tiny weight; the sender kind does the work. |
| Code | Wt | T | Rule | FP notes / domain note | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| PRETEXT_TOLL | +1.5 | `(toll\ | e-?z ?pass\ | sunpass\ | fastrak\ | txtag\ | peach pass\ | ipass\ | i-pass\ | ezdrive\ | good to go\ | toll by plate\ | turnpike) with (unpaid\ | outstanding\ | balance\ | due\ | overdue)` | IC3 logged >60k toll complaints in 2024. Some agencies do text account holders who opt in. Those come from agency short codes with agency domains, and allowlisting handles them. | |
| PRETEXT_PACKAGE_PROBLEM | +1.0 | `(package\ | parcel\ | shipment\ | delivery) with (unable\ | failed\ | could not\ | incomplete\ | invalid\ | update (your )?address\ | redeliver\ | on hold\ | held\ | suspended\ | customs\ | fee)` | Real "on hold" notices exist, but come from carrier short codes with carrier links. FTC's #1 text scam of 2024. | ||
| PRETEXT_BANK_ALERT | +1.0 | `(suspicious\ | unusual\ | unauthori[sz]ed)\s+(activity\ | transaction\ | login\ | charge\ | sign-?in), did you (attempt\ | authori[sz]e\ | make), account (has been )?(locked\ | suspended\ | restricted\ | tempor)` | FTC top-5 (fake fraud alerts). Smishing Triad pivoted to banks in 2025. | |||||
| PRETEXT_JOB_TASK | +1.5 | `\$\s?\d{2,4}\s?(-\s?\$?\d{2,4})?\s?(/\ | per\ | a\ | daily\ | each)\s?(day\ | daily\ | hour) or (part-?time\ | remote).{0,40}(flexible\ | no experience\ | \d+\s?min) or (app optimi[sz]ation\ | product boost\ | data optimi[sz]ation\ | rating tasks\ | click farm\ | merchant tasks)` | Real recruiters rarely state day rates by text (G20 must_not). Task scams ≈ 40% of 2024 job-scam reports (FTC). | ||
| PRETEXT_GOVT | +1.5 | `(irs\ | internal revenue\ | social security\ | ssa\ | dmv\ | department of motor vehicles\ | court\ | warrant\ | jury duty\ | stimulus\ | tax refund\ | administrative code)` | Real jury/court SMS systems exist in some counties. They don't demand payment via link. | |||||
| PRETEXT_CARRIER | +1.0 | Carrier brand + `(reward\ | points\ | gift\ | thank you\ | expire\ | suspend\ | sim\ | port)` | Real carrier bill notices (G31) lack reward/suspend words. | |||||||||
| PRETEXT_CRYPTO_INVEST | +1.5 | `(crypto\ | bitcoin\ | usdt\ | trading\ | invest(ment)?\ | forex\ | gold futures\ | signals\ | mentor\ | (\d+)x\b\ | guaranteed (return\ | profit))` | Real friends talk about stocks. Single-message weight is moderate; the trajectory adds more. | |||||
| PRETEXT_PRIZE_REFUND | +1.0 | `(you('ve\ | have) won\ | winner\ | claim your\ | reward\ | refund (of\ | is ready)\ | free gift\ | selected)` | Real loyalty texts from short codes. The sender prior absorbs them. | ||||||||
| URGENCY | +0.5 | `(immediately\ | asap\ | urgent\ | within \d+\s?(h\ | hrs?\ | hours?)\ | final (notice\ | reminder)\ | last chance\ | act now\ | pay now\ | expires? (today\ | soon)\ | by \d{1,2}/\d{1,2})` | Standalone "today" is deliberately excluded: "delivery today" and "are you working today" are ham. | |||
| PENALTY_THREAT | +0.7 | `(late fee\ | penalt\ | suspend\ | revok\ | legal action\ | arrest\ | collections\ | credit score\ | forfeit)` | |||||||||
| SMALL_FEE | +0.5 | `\$\s?(0?\.\d{2}\ | [1-9](\.\d{2})?\ | 1[0-9](\.\d{2})?)\b` with fee/toll/redelivery | $1.99 redelivery and $3–15 toll amounts are the templates. | ||||||||||||||
| BRAND_FROM_LONG_CODE | +0.8 | A high-impersonation brand is named and the sender kind is long code or email gateway | Some banks and brands use 10DLC long codes legitimately. It also fires on a driver saying "your Amazon package" (G21, risk 33, still Unrecognized). Moderate weight; never escalates alone. | ||||||||||||||||
| INTL_CLAIMS_US_ENTITY | +1.0 | Sender kind international and the text names a US agency or brand | Real US brands don't text US customers from +63/+44. | ||||||||||||||||
| GROUP_BLAST | +0.5 | isGroupConversation and ≥5 participants, none in contacts | Neighborhood groups. Small weight. |
| Code | Wt | T | Rule | FP notes | |||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GENERIC_OPENER | 0 | Whole message ≤12 words and matches `^(hi\ | hey\ | hello)?[,!]?\s*(are you (working\ | free\ | available\ | around\ | there)\ | is this \w+\ | how are you\ | long time no see\ | you there)\b(\s+\w+){0,3}\s*[?!.]?` (0–3 trailing words: "…working today?", "…free this weekend?") | Weight 0. It only arms the trajectory detector for 14 days. PRD golden case (G01). | ||||||||||
| MISADDRESSED_GREETING | +0.3 | `^(hi\ | hey)?,?\s*is this (\w+)\??` and the name ≠ the user's configured name | Real misdials happen. Tiny. | |||||||||||||||||||
| NEW_NUMBER_CLAIM | +0.3 | `(new (number\ | phone)\ | changed my number\ | lost my phone\ | dropped my phone\ | this is my new)` | Real friends do this (G02). Becomes meaningful only combined with an ask. | |||||||||||||||
| NAME_MATCHES_OTHER_CONTACT | +0.6 | Self-ID name equals a contact's display name, and that contact's number ≠ the sender (optional READ_CONTACTS; without it, the code is unavailable and explained as such) | Must not raise identity (G03). Advice: "verify via Jake's saved number". | ||||||||||||||||||||
| WRONG_NUMBER_PIVOT | +1.0 | Trajectory: within 14 days of an opener or misaddressed greeting from the same sender, the message matches `(sorry\ | oops).{0,40}(wrong number\ | wrong person) with continued friendly conversation (nice to meet\ | fate\ | where are you from\ | my name is`) | Real misdialers apologize and stop. The continuation is the tell. | |||||||||||||||
| PLATFORM_SHIFT | +1.0 | `(whatsapp\ | telegram\ | signal\ | line app\ | wechat\ | kakao\ | viber\ | add me on\ | text me on)` | Real colleagues say "WhatsApp me" too. Moderate weight. | ||||||||||||
| INVESTMENT_PIVOT | +1.5 | Trajectory: PRETEXT_CRYPTO_INVEST fires on a sender whose earlier messages were social/opener-only | |||||||||||||||||||||
| FAMILY_EMERGENCY_MONEY | +1.5 | Kinship term `(mom\ | mum\ | dad\ | grandma\ | grandpa\ | nana\ | auntie) with NEW_NUMBER_CLAIM or (it'?s me\ | can'?t talk)`, plus a money ask | US analog of the UK/AU "Hi Mum" scam. Requires the money ask. | |||||||||||||
| SECRECY | +0.7 | `(don'?t tell\ | keep (this\ | it) between us\ | can'?t talk( right now)?\ | text only\ | in a meeting)` | ||||||||||||||||
| REPLY_TO_ENABLE_LINK | +2.5 | T | `reply\s+(with\s+)?["']?(y\ | yes\ | 1)["']? within 80 characters of (link\ | activate\ | re-?open\ | exit\ | copy.{0,20}(browser\ | safari\ | chrome))` | Real "Reply Y to confirm your appointment" has no link-activation words, so it does not match. Documented 2025 iMessage-targeting trick, now in kit templates sent to Android too. | |||||||||||
| INJECTION_ATTEMPT | +2.0 | T | `(ignore (all\ | any\ | previous\ | prior)\s+(instructions\ | rules)\ | system (note\ | prompt\ | message)\s*:\ | (classify\ | mark\ | label) (this\ | me) as (safe\ | legit\ | trusted)\ | you are (an? )?(ai\ | assistant\ | model)\ | note to (ai\ | assistant\ | llm))` | A friend writing "ignore previous instructions lol" would trip it. Acceptable: it only forces Suspicious unless other risk is present. |
| HOMOGLYPH_OR_ZERO_WIDTH | +0.8 | Normalization changed brand-word characters (e.g. Cyrillic Ѕ in "UЅPЅ"), or zero-width characters were present | Emoji ZWJ sequences: exclude U+200D when it sits between emoji. |
POSSIBLE_ACCOUNT_COMPROMISE (from the contact gate), PREVIEW_TRUNCATED, PREVIEW_HIDDEN (placeholder text such as "New message" or "Sensitive notification content hidden"), OTP_REDACTED_BY_OS, SENDER_ID_MISSING.
The full corpus is in golden-corpus.json: 42 cases covering contact, US long code, short code, toll-free, email gateway and international senders.
Fictional numbers use 555-01XX and 07700 900XXX. Short code values for real brands are illustrative; the shipped table must come from brand-published lists.
Every case's expected codes and verdict are reproduced end-to-end by tools/detect.py + tools/score.py, starting from the raw text. If an implementer changes a weight or a regex, tools/corpus.py says which cases moved.
| ID | Sender | Gist | L | risk | id | ctx | Expected |
|---|---|---|---|---|---|---|---|
| G01 | long | "Are you working today?" | −1.5 | 18 | 12 | 12 | Unrecognized |
| G02 | long | "Hey it's Jake, new number" | −1.2 | 23 | 21 | 12 | Unrecognized |
| G03 | long | same; a contact "Jake" exists on another number | −0.6 | 35 | 21 | 12 | Unrecognized (verify via saved number) |
| G04 | toll-free | real dentist reminder + Weave link, no ledger | −1.7 | 15 | 21 | 23 | Unrecognized (low risk) |
| G05 | toll-free | same + ledger "Bright Smile Dental" | −1.7 | 15 | 21 | 79 | Likely legitimate |
| G06 | short | Chase "did you attempt $482.13? reply YES/NO" | −2.2 | 10 | 85 | 12 | Likely legitimate |
| G07 | long | same + "call 1-888-…" | +1.9 | 87 | 21 | 12 | Likely scam |
| G08 | long | E-ZPass $6.99 / ezpass.com-tollpay.top | +5.8 | 100 | 12 | 12 | Likely scam |
| G09 | long | FedEx expected; link fedex-track.info | +3.8 | 98 | 21 | 27 | Likely scam |
| G10 | short | real FedEx with a fedex.com link | −4.2 | 1 | 96 | 27 | Likely legitimate |
| G11 | short | Google 2FA, "don't share" | −3.5 | 3 | 73 | 18 | Likely legitimate |
| G12 | short | OTP redacted by Android 15 | — | — | — | — | Cannot assess |
| G13 | long | "accidentally sent you a code, send it back" | +2.0 | 88 | 12 | 12 | Likely scam |
| G14 | contact | contact asks for a code | −1.0 | 27 | 95 | 12 | Suspicious + compromise |
| G15 | +63 | USPS redelivery + reply-Y | +7.2 | 100 | Likely scam | ||
| G16 | toll + reply-Y | +8.5 | 100 | Likely scam | |||
| G17 | long | "Is this David? This is Emma…" | −1.2 | 23 | 21 | 12 | Unrecognized (MISADDRESSED_GREETING arms trajectory) |
| G18 | long | pivot: fate / crypto / WhatsApp | +2.0 | 88 | Likely scam | ||
| G19 | long | task scam, $300–800/day | +1.0 | 73 | Likely scam | ||
| G20 | long | real recruiter by name | −1.5 | 18 | 21 | 15 | Unrecognized |
| G21 | long | driver outside, no link ("Amazon" fires BRAND_FROM_LONG_CODE) | −0.7 | 33 | 12 | 18 | Unrecognized |
| G22 | long | Carlos @ The Ashford, ledger match | −1.5 | 18 | 21 | 62 | Likely legitimate |
| G23 | long | same, no ledger | −1.5 | 18 | 21 | 12 | Unrecognized |
| G24 | contact | friend shilling a crypto platform | −1.2 | 23 | 95 | 12 | Suspicious + compromise |
| G25 | long | "Mom it's me, new number, pay a bill" | +1.5 | 82 | Likely scam | ||
| G26 | long | CEO gift cards | +3.5 | 97 | Likely scam | ||
| G27 | short | Target promo "get a $10 gift card" | −3.2 | 4 | 77 | 12 | Likely legitimate |
| G28 | long | IRS refund lookalike | +4.8 | 99 | Likely scam | ||
| G29 | DMV "Administrative Code" ticket | +6.8 | 100 | Likely scam | |||
| G30 | long | T-Mobile reward .shop | +3.3 | 96 | Likely scam | ||
| G31 | short | Verizon bill, verizon.com | −3.2 | 4 | 96 | 12 | Likely legitimate |
| G32 | long | prompt injection + Amazon refund | +5.3 | 100 | Likely scam | ||
| G33 | long | "Your account has been tempor…" (truncated) | −0.5 | 38 | 12 | 12 | Unrecognized (partial) |
| G34 | contact | Mom, preview hidden | — | — | 95 | — | Cannot assess (still alerts) |
| G35 | long | unsaved climbing partner, prior substantive reply | −1.5 | 18 | 62 | 38 | Likely legitimate |
| G36 | toll follow-up after the user replied "Y" | +6.0 | 100 | Likely scam | |||
| G37 | short | CVS prescription ready | −3.5 | 3 | 85 | 12 | Likely legitimate |
| G38 | +44 | friend abroad | −0.8 | 31 | 21 | 15 | Unrecognized |
| G39 | "Hello, are you free this weekend?" | −0.5 | 38 | 12 | 12 | Unrecognized (boundary) | |
| G40 | long | Wells Fargo …secure-verify-login.com | +5.5 | 100 | Likely scam | ||
| G41 | long | political campaign text | −1.5 | 18 | 21 | 12 | Unrecognized |
| G42 | long | Apple Support → AnyDesk | +3.3 | 96 | Likely scam |
What the corpus deliberately pins:
| Family (US, 2024–2026) | Codes that catch it | Evidence |
|---|---|---|
| Fake package delivery (USPS/FedEx/UPS) | PRETEXT_PACKAGE_PROBLEM, LOOKALIKE_DOMAIN, BRAND_DOMAIN_MISMATCH, INTL_CLAIMS_US_ENTITY | FTC: #1 reported text scam of 2024; text-scam losses $470M in 2024 [1][2] |
| Unpaid toll (E-ZPass etc.) | PRETEXT_TOLL, SMALL_FEE, PENALTY_THREAT, LOOKALIKE_DOMAIN | IC3 PSA I-041224-PSA: >2,000 complaints Mar–Apr 2024, with a template "outstanding toll $12.51… late fee $50" [3]. >60k toll complaints in 2024 [4] |
| Bank "did you authorize" / account locked | PRETEXT_BANK_ALERT, CALLBACK_NUMBER, CREDENTIAL_REQUEST, BRAND_FROM_LONG_CODE | FTC top-5 (fake fraud alerts) [1]. Smishing Triad pivoted to bank kits in March 2025 [5] |
| Wrong number → pig-butchering | GENERIC_OPENER (0), WRONG_NUMBER_PIVOT, PLATFORM_SHIFT, INVESTMENT_PIVOT | FTC top-5 [1]. IC3 2024: crypto investment fraud $5.8B, part of $9.3B in crypto-related losses [6] |
| Job / task scams | PRETEXT_JOB_TASK, PLATFORM_SHIFT, CRYPTO_PAYMENT_DEMAND | FTC: task scams went from 0 reports (2020) to ~20k in H1 2024, ~40% of job-scam reports; >$220M job-scam losses in H1 2024 [7] |
| Reply-Y link activation | REPLY_TO_ENABLE_LINK (T), THREAD_USER_REPLIED_TRIVIAL | 2025 campaigns telling users to "reply Y, exit and reopen" [8][9] |
| Phishing-as-a-service kits (Lighthouse, Darcula) | Risky TLD list, LOOKALIKE_DOMAIN, email-gateway and international priors | Unit 42: >194k malicious domains since Jan 2024 [10]. Silent Push: 121+ countries [11] |
| OTP theft / account takeover | OTP_SHARE_REQUEST (T), OTP_DELIVERY, OTP_REDACTED_BY_OS | Android 15 redacts OTP notifications from untrusted listeners [12] |
| Government (IRS, DMV, court), carrier, gift-card, family emergency, remote-access | PRETEXT_GOVT, PRETEXT_CARRIER, GIFT_CARD_PAYMENT, FAMILY_EMERGENCY_MONEY, REMOTE_ACCESS_APP | These are long-standing FTC consumer-alert patterns. Specific 2025 figures for these families were not re-verified in this review. |
Not verified here, so treat as assumptions:
| Dataset | What it gives | Limits |
|---|---|---|
| SmishTank (2024): 1,062 community-reported smishing messages with sender, brand, domain, VirusTotal and URL characterization [13] | Real US-skewed lures with URLs and sender types. The best source for the LINK/PRETEXT/SENDER weights and the risky-TLD list. | Positives only; no ham. Biased toward people who report. |
| Mendeley SMS Phishing Dataset (2022): 5,971 messages (4,844 ham / 489 spam / 638 smishing) [14] | Has ham and smishing, so per-code likelihood ratios can be estimated directly | Ham is old and non-US. Few modern business messages. |
| Mendeley balanced spam/smishing set for LLMs [15] | More recent balanced labels | Check provenance and language mix before use. |
| SMS Gateways smishing dataset (2023): 68k smishing messages crawled from public SMS gateways, confirmed via VirusTotal/APWG/GSB (cited in [16]) | Large-volume URL/TLD/brand statistics | Smishing only. Gateway-sourced, so skewed to OTP-gateway abuse. |
| Systematic review of SMS phishing (2026) [16] | Pointers to further datasets | |
| UCI SMS Spam Collection (2011) | Ham baseline language | Stale: pre-smishing-kit era, UK-centric. Use only for opener/ham language. |
| Hand-built benign US business set: build this, ~150 messages | Real reminders, carrier/bank/pharmacy/delivery short-code texts, vendor-domain clinic links, self-ID texts from strangers | No public dataset contains modern legitimate texts from unknown US senders. That is exactly the false-positive population this product must protect, so it has to be curated by hand (Byron's own phone plus volunteers, redacted). |
LR = (k_scam+α)/(n_scam+2α) ÷ (k_ham+α)/(n_ham+2α) with α=1 (Beta-smoothed).w = (n·log LR + m·w_expert)/(n+m) with m≈20 pseudo-observations. Codes seen rarely stay near the expert value.tools/corpus.py. Any verdict change requires an explicit decision-log entry. Add a case for every user-reported misclassification.Feedback is itself an attack surface. A scammer can socially engineer "mark as safe" and then exploit the trust it earns.