Driver Drowsiness Detection System
End-to-end real-time driver fatigue monitoring system using:
MediaPipe Face Mesh for facial landmark extraction
a 14-dimensional temporal feature representation
BiLSTM + Transformer Encoders + Temporal Attention
FastAPI for real-time inference and authentication
Next.js for webcam monitoring and live alerts
WebSocket streaming with backpressure for low-latency inference
EMA smoothing + temporal voting for stable alert decisions
PyTorch + CUDA for training and GPU inference
The upgraded system predicts a binary driver state:
alert
drowsy
The original dataset still contains three behavioral categories:
alert
drowsy
microsleep
For Transformer training, the labels are mapped as:
alert -> 0
drowsy -> 1
microsleep -> 1
microsleep is therefore retained as a severe fatigue example rather than discarded.
The binary model output is converted into the user-facing alert levels:
safe
soft
warning
danger
through EMA smoothing, temporal voting and stateful transition logic.
- Current project status The project has moved beyond the original 10-feature BiLSTM pipeline. Current upgraded pipeline:
Webcam
|
v
Next.js LiveMonitor
|
| persistent WebSocket + backpressure
v
FastAPI
|
v
MediaPipe Face Mesh
|
v
9 absolute geometric features
|
v
~2-second timestamped temporal history
|
v
Pose stabilization + 60-step resampling
|
v
5 temporal delta features
|
v
60 x 14 feature sequence
|
v
Training normalization
|
v
LayerNorm
|
v
Dense 14 -> 64
|
v
Learnable Positional Encoding
|
v
BiLSTM (32 units/direction)
|
v
Transformer Encoder #1
|
v
Transformer Encoder #2
|
v
Temporal Attention Pooling
|
v
Dense(64) + GELU + Dropout
|
v
Dense(1)
|
v
Sigmoid
|
v
Drowsiness Probability
|
v
EMA
|
v
Temporal Voting
|
v
Stable ALERT / DROWSY
|
v
safe / soft / warning / danger
|
v
Audio Alarm
Verified milestones [x] 14-feature preprocessing [x] 60-step temporal representation [x] approximately 2-second time-normalized windows [x] head-pose stabilization [x] BiLSTM + Transformer architecture [x] video-level train / validation / test split [x] train-only normalization [x] balanced training sampler [x] GPU training [x] validation threshold search [x] deployment checkpoint [x] FastAPI checkpoint loading [x] live video -> feature -> Transformer smoke test [x] EMA + temporal voting decision engine [x] recovery behavior testing [x] Transformer WebSocket route [x] frontend WebSocket/backpressure design [ ] final Stage-6 live FPS / latency benchmarking [ ] larger subject-independent evaluation
- Key improvements over the original implementation
Property Original system Current upgraded system
Frame features 10 14
Temporal representation Fixed frame-count window ~2 s time-based window
Model input 45 x 10 / earlier variants 60 x 14
Temporal derivatives No Yes
Nose position No Yes
Sequence model BiLSTM + temporal attention BiLSTM + 2 Transformer encoders + temporal attention
Positional encoding No Learnable
Self-attention No Transformer 4 heads x 2 encoders
Model output 3-class softmax Binary drowsiness sigmoid
Runtime smoothing Heuristic fatigue score / hysteresis EMA + temporal voting + stable state logic
Streaming Periodic HTTP frames Persistent WebSocket + backpressure
Temporal FPS handling Frame-count dependent Timestamp-based resampling
Deployment model
drowsiness_bilstm.ptbest_transformer_deploy.ptMeasured test accuracy Older report ~81% 95.35% Test F1 Older report ~0.62 macro 0.9719 binary
The old and new metrics are not perfectly apples-to-apples because the old model used three output classes while the new model performs binary alert-vs-fatigue classification.
- Repository structure A representative current structure is:
driver-drowsiness-system/
|
|-- backend/
| |-- app/
| | |-- core/
| | | `-- config.py
| | |
| | |-- db/
| | |-- models/
| | |-- routes/
| | | |-- auth.py
| | | `-- inference.py
| | |
| | |-- schemas/
| | | `-- inference.py
| | |
| | `-- services/
| | |-- feature_extractor.py
| | |-- model_service.py
| | |-- session_state.py
| | |-- alert_engine.py
| | |
| | |-- transformer_model_arch.py
| | |-- transformer_model_service.py
| | |-- transformer_feature_extractor.py
| | |-- transformer_session_state.py
| | `-- transformer_decision_state.py
| |
| |-- test_transformer_model.py
| |-- test_transformer_runtime_pipeline.py
| `-- test_transformer_decision.py
|
|-- frontend/
| `-- src/
| |-- app/
| |-- components/
| | `-- LiveMonitor.tsx
| |-- lib/
| | `-- api.ts
| `-- types/
| `-- inference.ts
|
|-- ml/
| |-- datasets/
| | |-- raw/
| | | |-- alert/
| | | |-- drowsy/
| | | `-- microsleep/
| | |
| | `-- processed/
| | |-- X_transformer.npy
| | |-- y_transformer.npy
| | `-- meta_transformer.json
| |
| |-- checkpoints/
| | |-- best_transformer.pt
| | |-- best_transformer_deploy.pt
| | |-- transformer_feature_mean.npy
| | |-- transformer_feature_std.npy
| | |-- transformer_split.json
| | |-- transformer_metrics.json
| | `-- transformer_detailed_evaluation.json
| |
| `-- scripts/
| |-- feature_v2.py
| |-- preprocess_transformer.py
| |-- inspect_transformer_data.py
| |-- transformer_model.py
| |-- transformer_dataset.py
| |-- train_transformer.py
| `-- evaluate_transformer.py
|
`-- README.md
The original BiLSTM files may remain during migration, but the Transformer pipeline is the current architecture.
- High-level architecture
flowchart LR
CAM[Browser Webcam]
--> UI[Next.js LiveMonitor]
UI -->|JPEG frames over persistent WebSocket| API[
FastAPI
/api/v1/inference/ws/transformer/session_id
]
API --> MP[
MediaPipe Face Mesh
]
MP --> BASE[
9 Base Features
EAR L/R/Mean
MAR
Nose X/Y
Yaw/Pitch/Roll
]
BASE --> TEMP[
Timestamped ~2 s Session Buffer
]
TEMP --> RES[
Pose Stabilization
+ 60-Step Resampling
]
RES --> F14[
14-D Feature Sequence
+ Delta Features
]
F14 --> NORM[
Training Mean/Std Normalization
]
NORM --> MODEL[
BiLSTM
+ Transformer x2
+ Temporal Attention
]
MODEL --> PROB[
Drowsiness Probability
]
PROB --> EMA[
EMA Smoothing
]
EMA --> VOTE[
Temporal Voting
]
VOTE --> ALERT[
Stable Alert State
safe / soft / warning / danger
]
ALERT --> UI
AUTH[Auth Routes]
--> DB[(SQLite / PostgreSQL)]
- Frame-level feature representation The upgraded model uses 14 features per temporal step.
1 ear_left Left Eye Aspect Ratio
2 ear_right Right Eye Aspect Ratio
3 ear_mean Mean Eye Aspect Ratio
4 mar Mouth Aspect Ratio
5 nose_x Normalized nose X coordinate
6 nose_y Normalized nose Y coordinate
7 yaw Head yaw angle
8 pitch Head pitch angle
9 roll Head roll angle
10 delta_ear Change in mean EAR
11 delta_mar Change in MAR
12 delta_yaw Change in yaw
13 delta_pitch Change in pitch
14 delta_roll Change in roll
Therefore each temporal vector is:
f_t in R^14
and a complete model sample is:
X in R^(60 x 14)
- EAR and MAR 6.1 Eye Aspect Ratio For six eye landmarks:
EAR =
(||p2-p6|| + ||p3-p5||)
-----------------------
2 ||p1-p4||
The system calculates: left EAR right EAR mean EAR Lower sustained EAR generally corresponds to greater eye closure. 6.2 Mouth Aspect Ratio MAR represents mouth opening:
MAR =
(vertical lip distances)
------------------------
mouth width
- Head pose and angle stabilization
The backend estimates:
yaw
pitch
roll
using OpenCV
solvePnP. Euler angles can jump around the-180 / +180boundary. Example:
+179 deg -> -179 deg
Naive subtraction gives:
-358 deg
although the physical movement is approximately:
+2 deg
The current preprocessing/runtime pipeline therefore:
wraps angles consistently,
computes circular differences,
suppresses implausible source-frame pose jumps above approximately 45 deg,
calculates temporal pose deltas from a continuous pose representation.
This prevents large artificial delta_pitch or delta_roll spikes.
- Temporal window construction The model is designed around approximately 2 seconds of driver behavior. Instead of assuming every input is exactly 30 FPS, the system uses timestamp-aware resampling. Examples:
25 FPS source -> ~50 source observations over 2 s
30 FPS source -> ~60 source observations over 2 s
15 FPS source -> ~30 source observations over 2 s
All are transformed into:
60 temporal steps
x
14 features
- Dataset organization Place videos under:
ml/datasets/raw/
|
|-- alert/
|-- drowsy/
`-- microsleep/
Supported formats include:
.mp4
.avi
.mov
.mkv
.webm
Binary training label mapping:
alert -> 0
drowsy -> 1
microsleep -> 1
- Processed dataset Current processed Transformer dataset:
X shape: (6347, 60, 14)
y shape: (6347,)
dtype: float32
finite: True
Original-class sample distribution: Class Samples Alert 983 Drowsy 1029 Microsleep 4335 Total 6347 Binary distribution:
0 / alert = 983
1 / fatigue = 5364
- Data splitting strategy The project uses video-level splitting rather than random sequence-level splitting. This is important because adjacent temporal windows overlap heavily. Incorrect:
video A window 1 -> train
video A window 2 -> validation
video A window 3 -> test
Correct:
video A -> train only
video B -> validation only
video C -> test only
Actual split: Split Videos Samples Alert Positive Train 18 4554 720 3834 Validation 4 954 137 817 Test 4 839 126 713 Training original-class counts:
alert = 720
drowsy = 760
microsleep = 3074
- Feature normalization Normalization statistics are computed from the training split only:
x_norm = (x - mean_train) / std_train
The same mean/std values are reused for: validation test backend inference They are stored inside the model checkpoint and separately as:
ml/checkpoints/transformer_feature_mean.npy
ml/checkpoints/transformer_feature_std.npy
- Class balancing The training dataset is heavily biased toward microsleep windows. The training sampler therefore targets approximately:
alert -> 50% sampling mass
drowsy -> 25%
microsleep -> 25%
This produces approximate binary balance while preventing microsleep windows from dominating ordinary drowsiness. The loss therefore does not additionally use a positive-class weight.
- Transformer model architecture Model file:
ml/scripts/transformer_model.py
Architecture:
Input
(60, 14)
|
v
LayerNorm(14)
|
v
Linear
14 -> 64
|
v
Learnable Positional Encoding
(60 x 64)
|
v
BiLSTM
input = 64
hidden = 32
bidirectional = True
output = 64
|
v
Transformer Encoder #1
d_model = 64
heads = 4
FFN = 128
dropout = 0.10
|
v
Transformer Encoder #2
d_model = 64
heads = 4
FFN = 128
dropout = 0.10
|
v
Temporal Attention Pooling
|
v
Linear 64 -> 64
|
v
GELU
|
v
Dropout(0.30)
|
v
Linear 64 -> 1
|
v
Logit
|
v
Sigmoid during inference
|
v
Drowsiness Probability
Model size:
Total parameters: 105310
Trainable parameters: 105310
- Training configuration
Current training configuration:
Parameter Value
Batch size 64
Maximum epochs 40
Initial learning rate
3e-4Weight decay1e-4Optimizer AdamW Loss BCEWithLogitsLoss LR scheduler ReduceLROnPlateau Scheduler factor 0.5 Scheduler patience 2 Early stopping patience 7 Gradient clipping 1.0 Seed 42 Device CUDA when available Training GPU used NVIDIA GeForce RTX 4050 Laptop GPU The best validation checkpoint was obtained around epoch 7. Best validation loss:
0.09038
- Evaluation results 16.1 Refined validation threshold Initial training searched:
0.10 ... 0.90
and selected:
0.10
A later detailed validation search over:
0.01 ... 0.99
selected:
0.01
as the best validation-F1 threshold. Deployment checkpoint:
ml/checkpoints/best_transformer_deploy.pt
contains:
decision_threshold = 0.01
16.2 Validation metrics
At threshold 0.01:
Metric Value
Accuracy 99.06%
Precision 100.00%
Recall 98.90%
Specificity 100.00%
F1 0.9945
ROC-AUC 0.9999
PR-AUC 0.99998
Confusion matrix:
TN = 137
FP = 0
FN = 9
TP = 808
16.3 Test metrics Held-out test results: Metric Value Accuracy 95.35% Precision 99.85% Recall / Sensitivity 94.67% Specificity 99.21% F1 0.9719 ROC-AUC 0.9979 PR-AUC 0.9996 Test confusion matrix:
TN = 125
FP = 1
FN = 38
TP = 675
16.4 Original-class analysis Original Class Test Windows Mean Probability Predicted Drowsy Rate Alert 126 0.000864 0.79% Drowsy 86 0.998680 100.00% Microsleep 627 0.855435 93.94% The held-out drowsy video was detected correctly for all evaluated windows. Most remaining false negatives came from the held-out microsleep video.
The test split currently contains only four videos, so these results are promising but should not be interpreted as a complete real-world generalization study.
- Training commands From repository root:
python ml/scripts/preprocess_transformer.py
python ml/scripts/inspect_transformer_data.pyCompile model:
python -m py_compile ml/scripts/transformer_model.pyModel smoke test:
python ml/scripts/transformer_model.pyTrain:
python ml/scripts/train_transformer.pyDetailed evaluation:
python ml/scripts/evaluate_transformer.py- Generated ML artifacts Processed dataset:
ml/datasets/processed/
|-- X_transformer.npy
|-- y_transformer.npy
`-- meta_transformer.json
Checkpoint artifacts:
ml/checkpoints/
|-- best_transformer.pt
|-- best_transformer_deploy.pt
|-- transformer_feature_mean.npy
|-- transformer_feature_std.npy
|-- transformer_split.json
|-- transformer_metrics.json
`-- transformer_detailed_evaluation.json
best_transformer.pt
Original best training checkpoint.
best_transformer_deploy.pt
Deployment copy with the refined decision threshold:
0.01
- Backend Transformer services
transformer_model_arch.pyDeployment copy of the PyTorch architecture. It must match the training architecture exactly sostate_dictloading is valid.transformer_model_service.pyResponsibilities: resolve checkpoint path load checkpoint load architecture config load trained normalization statistics normalize(60, 14)sequences run GPU/CPU inference convert logit through sigmoid apply decision threshold The service intentionally does not silently fall back to heuristic inference. If loading fails, the Transformer error is surfaced explicitly.transformer_feature_extractor.pyExtracts the nine runtime base features:
EAR left
EAR right
EAR mean
MAR
nose X
nose Y
yaw
pitch
roll
transformer_session_state.py
Responsibilities:
timestamp incoming observations
keep approximately two seconds of source history
stabilize head pose
resample to 60 temporal steps
compute five delta features
return a finite (60, 14) sequence
Current important constants:
SEQ_LEN = 60
WINDOW_SECONDS = 2.0
MIN_SOURCE_SAMPLES = 12
MAX_POSE_STEP_DEGREES = 45.0
transformer_decision_state.py
Responsibilities:
EMA probability smoothing
recent prediction voting
stable alert/drowsy state
recovery logic
UI alert severity
- EMA + temporal voting The model's raw probability is not used to trigger the alarm directly. 20.1 EMA Conceptually:
EMA_t =
alpha * probability_t
+
(1-alpha) * EMA_(t-1)
Current asymmetric behavior:
rise alpha = 0.50
fall alpha = 0.70
The larger fall alpha helps the detector recover faster when the driver becomes alert again. 20.2 Voting Recent predictions are stored in a five-element vote window. Example:
0 0 1 1 1
gives:
vote ratio = 3 / 5 = 0.60
Current decision settings:
vote window = 5
minimum votes = 3
enter ratio = 0.60
exit ratio = 0.20
The ML model remains binary:
alert
drowsy
The operational severity is separately mapped to:
safe
soft
warning
danger
- Decision-engine verified behavior Decision smoke testing produced the intended behavior. Stable alert Low probabilities remain:
prediction = alert
level = safe
Drowsiness onset After sustained high probabilities:
safe
-> soft
-> warning
-> danger
Recovery After the probability falls:
danger
-> warning
-> safe
The recovery path was intentionally designed to avoid the old behavior where the alarm could continue long after the driver woke up.
- Backend API reference Base prefix:
/api/v1
22.1 Health
GET /healthExample:
{
"status": "ok"
}22.2 Auth Existing auth endpoints remain:
POST /api/v1/auth/register
POST /api/v1/auth/login
GET /api/v1/auth/me
22.3 Legacy HTTP inference During migration the original route may remain available:
POST /api/v1/inference/frame
This is the old BiLSTM/hybrid path and is not the preferred Transformer streaming path. 22.4 Transformer HTTP inference Parallel Transformer endpoint:
POST /api/v1/inference/transformer/frame
Request:
{
"session_id": "demo-session",
"frame_base64": "data:image/jpeg;base64,..."
}Possible statuses:
no_face_detected
collecting
ok
error
22.5 Transformer WebSocket Current real-time Transformer route:
/api/v1/inference/ws/transformer/{session_id}
Example local URL:
ws://127.0.0.1:8000/api/v1/inference/ws/transformer/demo-session
Client message:
{
"frame_base64": "data:image/jpeg;base64,..."
}Example collecting response:
{
"status": "collecting",
"session_id": "demo-session",
"sequence_length": 42,
"sequence_ready": false,
"source_samples": 19,
"coverage_seconds": 1.31,
"message": "Collecting approximately 2 seconds of temporal face features.",
"source": "transformer"
}Example ready response:
{
"status": "ok",
"session_id": "demo-session",
"sequence_length": 60,
"sequence_ready": true,
"source_samples": 28,
"coverage_seconds": 2.01,
"score": 93.1,
"level": "danger",
"prediction": "drowsy",
"message": "Critical drowsiness alert. Wake up and stop safely.",
"source": "transformer",
"raw_probability": 0.9982,
"smoothed_probability": 0.931,
"decision_threshold": 0.01,
"vote_ratio": 0.8
}- Frontend real-time behavior
LiveMonitor.tsxuses the Transformer WebSocket instead of low-rate HTTP polling. Main responsibilities: open webcam show live feed downscale inference frames JPEG-compress frames open Transformer WebSocket apply backpressure update inference status display temporal progress display raw probability display EMA probability display vote ratio display prediction play progressive alarm patterns optionally record webcam video Recommended/current inference-frame configuration:
inference width = 640 px
target send interval ~= 75 ms
This gives an upper target near:
13.3 FPS
- Why WebSocket backpressure matters The frontend sends a new frame only when the previous frame has been processed. Conceptually:
send frame
|
wait for inference
|
receive response
|
send next frame
rather than:
send
send
send
send
send
...
This prevents inference queues from accumulating stale frames. That is especially important when the driver transitions from:
drowsy -> alert
because an old frame queue could otherwise keep the alarm active even after the driver's current state has changed.
- Runtime smoke tests
25.1 Checkpoint loading
From
backend/:
python test_transformer_model.pyVerified output includes:
Loaded: True
Sequence length: 60
Input dimension: 14
Threshold: 0.01
Example verified predictions:
alert sample:
probability ~= 0.00048
prediction = alert
drowsy sample:
probability ~= 0.99873
prediction = drowsy
25.2 Full runtime video pipeline
python test_transformer_runtime_pipeline.pyVerified alert video example:
Sequence shape: (60, 14)
Finite: True
Prediction: alert
Probability ~= 0.00048
Verified drowsy video example:
Sequence shape: (60, 14)
Finite: True
Prediction: drowsy
Probability ~= 0.99874
25.3 Decision logic
python test_transformer_decision.py- Configuration Backend settings are read from:
backend/app/core/config.py
Important model paths:
legacy model path:
../storage/models/drowsiness_bilstm.pt
Transformer deployment path:
../ml/checkpoints/best_transformer_deploy.pt
Typical frontend environment:
NEXT_PUBLIC_API_BASE_URL=http://127.0.0.1:8000/api/v1- Local development setup 27.1 Prerequisites Recommended: Python 3.11 Node.js 20+ npm webcam access Windows / Linux / macOS optional CUDA-compatible GPU The current Transformer was trained and tested on:
NVIDIA GeForce RTX 4050 Laptop GPU
- Backend setup From repository root:
cd backend
python -m venv .venvWindows PowerShell:
.\.venv\Scripts\Activate.ps1Linux/macOS:
source .venv/bin/activateInstall dependencies:
pip install -r requirements.txtRun:
uvicorn app.main:app --reload --host 127.0.0.1 --port 8000Useful endpoints:
http://127.0.0.1:8000/health
http://127.0.0.1:8000/docs
- Frontend setup From repository root:
cd frontend
npm installCreate/update:
frontend/.env.local
with:
NEXT_PUBLIC_API_BASE_URL=http://127.0.0.1:8000/api/v1Build:
npm run buildRun development server:
npm run devOpen:
http://localhost:3000
Monitor page:
http://localhost:3000/monitor
- End-to-end startup Terminal 1:
cd backend
.\.venv\Scripts\Activate.ps1
uvicorn app.main:app --reloadTerminal 2:
cd frontend
npm run devThen open:
http://localhost:3000/monitor
Click:
Start Camera
Initial state:
status = collecting
sequence < 60
coverage < ~2 seconds
When enough history exists:
status = ok
sequence = 60 / 60
source = transformer
- Frontend diagnostic fields The monitoring interface should expose:
Session
Backend status
Sequence progress
Temporal coverage
Source frame count
Raw probability
EMA probability
Vote ratio
Score
Prediction
Alert level
Inference source
- Database model
Existing SQLAlchemy models include:
usersdriving_sessionsalert_eventsRelationships:
User 1..N DrivingSession
DrivingSession 1..N AlertEvent
The current ML upgrade primarily changes the inference stack; authentication and persistence modules remain structurally separate.
- Legacy path During development/migration the repository may still contain:
feature_extractor.py
model_service.py
session_state.py
alert_engine.py
preprocess_videos.py
model.py
train.py
These belong to the earlier BiLSTM pipeline.
New Transformer production files use the transformer_* names to avoid breaking the previous implementation during migration.
Once the Transformer frontend and Stage-6 live testing are complete, the legacy path can be retired or moved under a legacy/ directory.
- Troubleshooting 34.1 Transformer model is not loading Run:
cd backend
python test_transformer_model.pyCheck:
Loaded: True
Load error: None
Verify:
ml/checkpoints/best_transformer_deploy.pt
34.2 Always seeing collecting
Check:
webcam face is visible
temporal coverage is increasing
source frame count is increasing
WebSocket is connected
the frontend is not using the old low-rate HTTP route
Runtime sequence readiness requires approximately:
1.85 - 2.0 seconds
34.3 no_face_detected
Possible causes:
face outside frame
poor lighting
camera angle too extreme
strong occlusion
sunglasses / landmark failure
The Transformer temporal context is reset when the face disappears.
34.4 Alarm remains active too long Inspect:
raw_probability
smoothed_probability
vote_ratio
level
The current decision engine uses faster EMA decay when probability falls. Also confirm WebSocket backpressure is active so old frames are not queued.
34.5 Invalid checkpoint state dict The backend architecture must match:
ml/scripts/transformer_model.py
exactly. Do not rename internal PyTorch module fields unless the checkpoint is retrained or migrated.
34.6 CUDA is unavailable Check:
python -c "import torch; print(torch.cuda.is_available())"34.7 Frontend cannot connect to Transformer WebSocket Check backend:
http://127.0.0.1:8000/health
Check environment:
NEXT_PUBLIC_API_BASE_URL=http://127.0.0.1:8000/api/v1Expected WebSocket URL:
ws://127.0.0.1:8000/api/v1/inference/ws/transformer/demo-session
- Stage-6 real-time validation Offline Transformer evaluation is complete. The final deployment-validation stage should measure: Metric Status Effective processed FPS Pending final measurement Mean frame round-trip latency Pending final measurement P95 inference latency Pending final measurement Drowsiness alert onset delay Pending final measurement Wake-up recovery delay Pending final measurement Long-run WebSocket stability Pending final measurement Do not publish guessed values. Recommended Stage-6 tests: 30-60 seconds normal alert behavior. Sustained eye closure. Repeated slow blinking. Head-drop simulation. Yawn-like mouth behavior. Recovery after danger state. Temporary face disappearance. Low-light behavior. High CPU/GPU load. 5-10 minute continuous WebSocket run.
- Known limitations
Current limitations include:
The test set contains only four held-out videos.
Binary training combines drowsy and microsleep into one positive class.
Raw sigmoid scores are not perfectly probability-calibrated.
Current best validation threshold is unusually low (
0.01). Landmark performance can degrade under strong occlusion or poor lighting. Subject-level train/test separation should be used when reliable subject IDs are available. Webcam-based fatigue estimation is not a medical diagnostic system. Final Stage-6 FPS and latency benchmarks still need to be measured on the completed frontend.
- Production hardening
Before production deployment, consider:
Move JWT/secret keys to a secure secret manager.
Use Authorization header-based authentication.
Persist driving sessions and alert events during inference.
Add Alembic migrations instead of relying only on
create_all. Add WebSocket authentication. Add maximum JPEG/request-size checks. Add connection-rate and inference-rate limiting. Add structured logs for: inference latency MediaPipe failures model errors WebSocket reconnects Add session cleanup / expiry. Add Prometheus/OpenTelemetry metrics. Add calibrated decision-threshold evaluation on a larger dataset. Evaluate subject-independent generalization. Consider ONNX/TensorRT for lower latency.
- Useful commands cheat sheet Backend
cd backend
uvicorn app.main:app --reloadFrontend
cd frontend
npm run devTransformer preprocessing
python ml/scripts/preprocess_transformer.py
python ml/scripts/inspect_transformer_data.pyModel test
python ml/scripts/transformer_model.pyTrain
python ml/scripts/train_transformer.pyDetailed evaluation
python ml/scripts/evaluate_transformer.pyBackend checkpoint test
cd backend
python test_transformer_model.pyRuntime video pipeline test
cd backend
python test_transformer_runtime_pipeline.pyDecision logic test
cd backend
python test_transformer_decision.py- Fast file index
Backend
backend/app/main.pybackend/app/core/config.pybackend/app/routes/inference.pybackend/app/schemas/inference.pybackend/app/services/transformer_model_arch.pybackend/app/services/transformer_model_service.pybackend/app/services/transformer_feature_extractor.pybackend/app/services/transformer_session_state.pybackend/app/services/transformer_decision_state.pyFrontendfrontend/src/components/LiveMonitor.tsxfrontend/src/lib/api.tsfrontend/src/types/inference.tsMachine Learningml/scripts/feature_v2.pyml/scripts/preprocess_transformer.pyml/scripts/inspect_transformer_data.pyml/scripts/transformer_model.pyml/scripts/transformer_dataset.pyml/scripts/train_transformer.pyml/scripts/evaluate_transformer.pyCheckpointsml/checkpoints/best_transformer.ptml/checkpoints/best_transformer_deploy.ptml/checkpoints/transformer_feature_mean.npyml/checkpoints/transformer_feature_std.npyml/checkpoints/transformer_metrics.jsonml/checkpoints/transformer_detailed_evaluation.json
- Current model summary
Task:
Binary driver fatigue detection
Input:
~2 seconds of facial behavior
Input tensor:
60 x 14
Architecture:
LayerNorm
-> Linear(14,64)
-> Learnable Position Encoding
-> BiLSTM(32 bidirectional)
-> Transformer Encoder x2
-> Temporal Attention Pooling
-> Linear(64,64)
-> GELU
-> Dropout(0.30)
-> Linear(64,1)
Parameters:
105,310
Optimizer:
AdamW
Loss:
BCEWithLogitsLoss
Deployment threshold:
0.01
Test:
Accuracy 95.35%
Precision 99.85%
Recall 94.67%
Specificity 99.21%
F1 0.9719
ROC-AUC 0.9979
PR-AUC 0.9996
Runtime decision:
Sigmoid probability
-> EMA
-> temporal voting
-> stable alert/drowsy state
Deployment:
Next.js
-> WebSocket
-> FastAPI
-> MediaPipe
-> PyTorch Transformer
-> live alert UI
- Future work Potential next improvements: explicit subject-level splitting larger and more diverse test set probability calibration / temperature scaling separate drowsy-vs-microsleep severity head infrared / night-driving support personalized EAR/MAR baselines ONNX / TensorRT export edge-device deployment multimodal fusion with vehicle telemetry improved head-pose estimation model quantization continuous real-world driver evaluation
- Important note on reported metrics The current offline metrics are based on the current video-level split and should be reported with the evaluation protocol. They should not be presented as universal real-world performance. Final real-time FPS, latency, alert-onset time and recovery time must be measured during Stage 6 before being added to project documentation or reports.