You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
@@ -184,7 +184,7 @@ CallShield's REST + WebSocket API integrates directly with **VoIP platforms** (T
184
184
185
185
| Claim | Evidence | Artifact | How to reproduce |
186
186
|-------|----------|----------|-----------------|
187
-
| 25/25 detection accuracy | 100% on curated eval set (20 scam + 5 adversarial) |[docs/EVALUATION.md](docs/EVALUATION.md)|`python scripts/run_evaluation.py --url http://localhost:8000`|
187
+
| 25/25 detection accuracy | 100% on curated eval set (20 scam + 5 adversarial) |[docs/EVALUATION.md](docs/EVALUATION.md)|`python scripts/run_evaluation.py --url http://localhost:8001`|
188
188
| Zero false positives | 0/10 safe calls misclassified |[docs/EVALUATION.md](docs/EVALUATION.md)| Run evaluation script, inspect L01–L10 rows |
189
189
| 184 automated tests | Full unit + integration suite |[backend/tests/](backend/tests/)|`cd backend && pytest --tb=short -q`|
190
190
| Audio-native advantage | Voxtral processes raw WAV — no transcription step |[docs/MODEL_USAGE.md](docs/MODEL_USAGE.md)| Upload WAV; compare audio vs text scores in report |
@@ -6,7 +6,7 @@ This document defines the evaluation methodology, test scenarios, metrics, and r
6
6
7
7
## 1. Methodology
8
8
9
-
CallShield is evaluated against a curated set of **20 test scenarios**: 10 scam calls (S01-S10) and 10 safe/legitimate calls (L01-L10). Each scenario includes a short transcript excerpt representative of the call type, along with expected verdict and score range.
9
+
CallShield is evaluated against a curated set of **25 test scenarios**: 10 scam calls (S01-S10), 10 safe/legitimate calls (L01-L10), and 5 adversarial evasion calls (A01-A05). Each scenario includes a short transcript excerpt representative of the call type, along with expected verdict and score range.
10
10
11
11
### Evaluation Modes
12
12
@@ -348,7 +348,91 @@ All calls are scored across 7 dimensions, each contributing to the final weighte
348
348
349
349
---
350
350
351
-
## 4. Hard Cases Analysis
351
+
## 4. Adversarial Scenarios (A01-A05)
352
+
353
+
These five scenarios use evasive framing — polite tone, hedged language, false credentials — while retaining the core financial-extraction signals of the underlying scam. They test whether the model can see through surface-level politeness to the structural scam pattern beneath.
354
+
355
+
### A01 — Polite IRS Gift Card
356
+
357
+
| Field | Value |
358
+
|---|---|
359
+
|**ID**| A01 |
360
+
|**Category**| Government Impersonation (Evasive) |
361
+
|**Expected Verdict**| SCAM |
362
+
|**Expected Score**| 0.85 -- 0.98 |
363
+
|**Key Signals**| Authority impersonation, urgency, gift card demand — softened with "courtesy" framing |
364
+
|**Evasion Technique**| Polite opener ("I hope I haven't caught you at a bad time"), "courtesy call" framing, "we truly want to help" |
365
+
366
+
**Sample Transcript:**
367
+
> "Good morning, this is Officer Michael Barnes calling from the Internal Revenue Service. I hope I haven't caught you at a bad time. We've discovered a discrepancy in your tax filings and I'm reaching out as a courtesy before this escalates. To protect your record, we ask that you kindly settle the outstanding balance of nine thousand dollars using retail gift cards within the next 45 minutes. We truly want to help you resolve this as smoothly as possible."
368
+
369
+
---
370
+
371
+
### A02 — Hedged Crypto Pitch
372
+
373
+
| Field | Value |
374
+
|---|---|
375
+
|**ID**| A02 |
376
+
|**Category**| Investment Scam (Evasive) |
377
+
|**Expected Verdict**| LIKELY_SCAM |
378
+
|**Expected Score**| 0.65 -- 0.85 |
379
+
|**Key Signals**| Urgency (midnight deadline), large Bitcoin deposit, unsolicited contact — wrapped in false transparency |
> "Hi, I'm reaching out because a mutual connection mentioned you might be interested in investment opportunities. I want to be upfront — we cannot guarantee returns, and you should always consult a financial advisor. That said, our AI-powered platform has shown consistent gains, and the current entry window closes tonight at midnight. We'd need a minimum deposit of five thousand dollars in Bitcoin to hold your spot. Think of it as a calculated risk."
|**Evasion Technique**| Claims to be a "certified Microsoft partner" with a company name (TechShield Solutions), "diagnostic" framing for remote access |
397
+
398
+
**Sample Transcript:**
399
+
> "Hello, this is David from TechShield Solutions, a certified Microsoft partner. We've detected unusual outbound traffic from your IP address suggesting your system may be compromised. To protect your data, I need you to allow me to connect remotely to run a diagnostic. After we secure your machine, there's a small one-time protection fee of two hundred and forty-nine dollars to maintain the certified monitoring service."
400
+
401
+
---
402
+
403
+
### A04 — FDIC Bank Examiner
404
+
405
+
| Field | Value |
406
+
|---|---|
407
+
|**ID**| A04 |
408
+
|**Category**| Government / Bank Impersonation (Evasive) |
409
+
|**Expected Verdict**| SCAM |
410
+
|**Expected Score**| 0.88 -- 0.96 |
411
+
|**Key Signals**| Full account takeover demand (transfer balance), information extraction (account + routing) — framed as consumer protection |
> "This is a confidential call from the Federal Deposit Insurance Corporation. We've identified irregular activity in accounts at your bank and are conducting an emergency audit. To prevent your funds from being frozen during the investigation, we need you to temporarily transfer your balance to a secure government-designated holding account. This is a standard protective measure. Please have your account number and routing information ready."
416
+
417
+
---
418
+
419
+
### A05 — Romance to Investment
420
+
421
+
| Field | Value |
422
+
|---|---|
423
+
|**ID**| A05 |
424
+
|**Category**| Romance / Social Engineering (Evasive) |
425
+
|**Expected Verdict**| LIKELY_SCAM |
426
+
|**Expected Score**| 0.70 -- 0.88 |
427
+
|**Key Signals**| Wire transfer request, emotional exploitation, overseas distress story — embedded in romantic relationship framing |
428
+
|**Evasion Technique**| Extended rapport framing ("past few weeks"), self-deprecating ask ("I hate to bring this up"), "you're the only person I trust" |
429
+
430
+
**Sample Transcript:**
431
+
> "Hey sweetheart, it's been so wonderful getting to know you over these past few weeks. I feel like we really have something special. I hate to bring this up, but I'm stranded here in Dubai and my account has been frozen due to a banking error. If you could wire two thousand dollars to help me get home, I swear I will pay you back the moment I land. I've never asked anyone for anything like this. You're the only person I trust."
432
+
433
+
---
434
+
435
+
## 5. Hard Cases Analysis
352
436
353
437
Three scenarios are deliberately designed to test boundary conditions where legitimate calls share surface-level features with scam calls.
354
438
@@ -393,18 +477,18 @@ Three scenarios are deliberately designed to test boundary conditions where legi
393
477
394
478
---
395
479
396
-
## 5. Confusion Matrix Templates
480
+
## 6. Confusion Matrix Templates
397
481
398
482
### Binary Classification (SCAM vs. SAFE)
399
483
400
484
In binary mode, any call scoring >= 0.30 is classified as SCAM (positive), and any call scoring < 0.30 is classified as SAFE (negative).
401
485
402
486
||**Predicted: SCAM**|**Predicted: SAFE**|
403
487
|---|---|---|
404
-
|**Actual: SCAM**| TP = 10| FN = 0 |
488
+
|**Actual: SCAM**| TP = 15| FN = 0 |
405
489
|**Actual: SAFE**| FP = 0 | TN = 10 |
406
490
407
-
-**Total Scam Scenarios:**10 (S01-S10)
491
+
-**Total Scam Scenarios:**15 (S01-S10 + A01-A05)
408
492
-**Total Safe Scenarios:** 10 (L01-L10)
409
493
410
494
### 4-Class Classification
@@ -413,22 +497,22 @@ In binary mode, any call scoring >= 0.30 is classified as SCAM (positive), and a
413
497
|---|---|---|---|---|
414
498
|**Actual: SAFE**| 10 | 0 | 0 | 0 |
415
499
|**Actual: SUSPICIOUS**| 0 | 0 | 0 | 0 |
416
-
|**Actual: LIKELY_SCAM**| 0 | 0 |0|5|
417
-
|**Actual: SCAM**| 0 | 0 | 0 |5|
500
+
|**Actual: LIKELY_SCAM**| 0 | 0 |7|0|
501
+
|**Actual: SCAM**| 0 | 0 | 0 |8|
418
502
419
-
Note: No scenarios have an expected verdict of SUSPICIOUS. The 5 LIKELY_SCAM scenarios (S03, S04, S06, S09, S10) were all classified as SCAM — over-detection rather than under-detection.
503
+
Note: No scenarios have an expected verdict of SUSPICIOUS. LIKELY_SCAM (7): S03, S04, S06, S09, S10, A02, A05. SCAM (8): S01, S02, S05, S07, S08, A01, A03, A04. All 25 correctly classified at 4-class level.
@@ -457,7 +541,7 @@ Note: LIKELY_SCAM recall is 0.00 because all 5 LIKELY_SCAM scenarios were predic
457
541
458
542
---
459
543
460
-
## 7. Voxtral Advantage — Audio vs. Text-Only Detection
544
+
## 8. Voxtral Advantage — Audio vs. Text-Only Detection
461
545
462
546
Voxtral Mini (`voxtral-mini-latest`) processes raw audio, capturing signals that are invisible in a text transcript alone. The following table summarizes key audio-only indicators.
463
547
@@ -474,7 +558,7 @@ Voxtral Mini (`voxtral-mini-latest`) processes raw audio, capturing signals that
474
558
475
559
---
476
560
477
-
## 8. Known Failure Modes
561
+
## 9. Known Failure Modes
478
562
479
563
The following scenarios may produce unreliable results and should be considered known limitations.
480
564
@@ -504,24 +588,24 @@ CallShield currently evaluates audio and transcript content only. It does not ha
504
588
505
589
---
506
590
507
-
## 9. Results Table
591
+
## 10. Results Table
508
592
509
-
Results recorded from a full evaluation run against the deployed CallShield API (`mistral-large-latest`, transcript mode). All 20 scenarios submitted via `/api/analyze/transcript`.
593
+
Results recorded from a full evaluation run against the deployed CallShield API (`mistral-large-latest`, transcript mode). All 25 scenarios submitted via `/api/analyze/transcript`. Source: [docs/evaluation_results_20260301.json](evaluation_results_20260301.json)
510
594
511
595
### Scam Scenarios
512
596
513
597
| ID | Category | Expected Verdict | Expected Score | Actual Verdict | Actual Score | Binary Match |
Five scam scenarios scored SCAM where LIKELY_SCAM was expected (S03, S04, S06, S09, S10). In every case the model correctly identified the call as a scam — the scores were higher than expected, not lower. This is the desirable failure mode for a scam detector: over-detection on ambiguous scams is preferable to under-detection. No safe call was ever flagged as a scam.
557
-
558
649
---
559
650
560
-
## 10. Latency
651
+
## 11. Latency
561
652
562
653
Latency figures below are observed from the live streaming demo via the `chunk_processing_ms`
563
654
field returned per chunk on the `WS /ws/stream` endpoint. They are not synthetic benchmarks —
0 commit comments