Serve AstaBrief 8B (Allen AI) on Runpod Serverless with vLLM. License: apache-2.0.
| Setting | Value |
|---|---|
| Engine | vllm |
| Image | runpod/worker-v1-vllm:v2.27.0 |
| GPU | 1x RTX 4090 |
| Precision | bf16 |
| Max model length | 8192 |
| Variable | Value |
|---|---|
GPU_MEMORY_UTILIZATION |
0.90 |
MAX_CONCURRENCY |
30 |
MAX_MODEL_LEN |
8192 |
MODEL_NAME |
allenai/AstaBrief_8B |
TENSOR_PARALLEL_SIZE |
1 |
| GPU | $/hr | Startup s | TTFT ms | tok/s | $/1M output tokens |
|---|---|---|---|---|---|
| 1x RTX 4090 bf16 | $0.69 | 242.3 | 216.1 | 56.2 | $3.41 |
| 1x RTX 4090 fp8 | $0.69 | 349.3 | 210.3 | 77.6 | $2.47 |
Cost per 1M output tokens is the hourly rate divided by measured throughput. It assumes one saturated worker and no idle time, so treat it as a floor.
- Deploy on Runpod: https://console.runpod.io/deploy?template=c1bout9ej5&utm_source=hub&utm_medium=product&utm_campaign=202610_activation_indie-ml-dev_allen-ai-astabrief-8b&utm_content=readme
- Model page: https://www.runpod.io/models/allen-ai-astabrief-8b?utm_source=hub&utm_medium=product&utm_campaign=202610_activation_indie-ml-dev_allen-ai-astabrief-8b&utm_content=readme
- Docs: https://docs.runpod.io/public-endpoints/models/allen-ai-astabrief-8b?utm_source=hub&utm_medium=product&utm_campaign=202610_activation_indie-ml-dev_allen-ai-astabrief-8b&utm_content=readme
Every link carries utm_campaign=202610_activation_indie-ml-dev_allen-ai-astabrief-8b. Keep it intact when you copy a link anywhere else.
Run the suite against a live endpoint:
node hub-test-suite.mjs --repo <owner>/<name> --prefix allen-ai-astabrief-8b- --createrunpodctl serverless create --hub-id <vllm listing> --model-reference hf://allenai/AstaBrief_8B