GPU inference and benchmarking
Model serving, resource allocation, and performance testing for GPU-backed services.
WHAT CHANGED
- Repartitioned an 8-GPU vision-language plane across four model configurations using measured memory and throughput behavior.
- Corrected a benchmark that overstated throughput by roughly 4.6× when repeated inputs hit caches.
- Built a GPU vector pipeline recorded at about 33,480 items an hour with both GPUs fully utilized.
TECHNOLOGIES