Private VPC deployments, 3B–8B parameter open-weights, edge inference, and task-specific model arbitrage.