The bank’s infrastructure team successfully integrated nearly 10,000 heterogeneous accelerator cards into a single control plane. By deploying a stack that includes Kubernetes, Kueue, KEDA, Prometheus, HAMi, and Fluid, the organization brought 99% of its compute resources under a unified management framework. This shift increased average accelerator utilization from 35% to over 60% and reduced the cost of processing one million tokens by more than 60%.
Beyond hardware pooling, the bank introduced its proprietary Twinkle training framework to optimize multi-tenant fine-tuning. By allowing five LoRA tenants to share a single base model instance, the system cut accelerator resource usage by 80% and boosted training density fivefold. PeiXiang Tan, the bank's AI Infrastructure Architect, noted that the platform enables training and inference to follow workload-specific paths while maintaining end-to-end observability. Future development will focus on serverless inference scaling and dynamic concurrency management to further refine capacity utilization.




Comments (0)
No comments yet. Be the first!