openqareer

AI Architecture & Infrastructure Engineer

CMB Wing Lung Bank Limited · Hong Kong

# AI Architecture & Infrastructure Engineer **CMB Wing Lung Bank Limited** · Hong Kong · `On-site` 🕒 **Статус:** *Опубликовано: 4 дня назад* · *Источник: Indeed* --- ### About the Role CMB Wing Lung Bank Limited is a Licensed Bank registered in Hong Kong, which is formerly known as Wing Lung Bank. Founded on 25th February 1933, CMB Wing Lung Bank is among the oldest local Chinese banks in Hong Kong. In 2008, CMB Wing Lung Bank was acquired by China Merchants Bank (CMB). In 2009, the Bank has become a wholly-owned subsidiary of CMB. In 2018, the Bank has renamed as CMB Wing Lung Bank. CMB Wing Lung Bank and its subsidiaries hold multi licenses in banking, insurance, securities, trust and asset management sectors. Serving customers with whole-hearted, the Bank provides comprehensive banking products and services including retail banking, corporate banking and financial markets, etc. Among them, private banking, syndicated loans, bond issuance, financial markets, asset custody and asset management are influential and competitive in the market. At present, the Bank has over 30 branches and business outlets located in mainland China, Hong Kong, Macau and overseas, etc regions. The Bank has nearly 2,000 employees. Following the motto of “Progress with prudence, service with sincerity”, the Bank embraces “Serving with Heart, Quest for efficiency, Caring for sincerity” service value into the daily operation. The Bank strives to build a first-class, innovative and cutting-edge commercial bank in Hong Kong and provides an integrated local and cross-border financial and banking services for customers. Position Overview We are hiring an AI Architecture & Infrastructure Engineer to design and build bank-grade AI infrastructure, model platforms, and intelligent-application foundations. The role is responsible for turning AI capabilities into secure, reliable, scalable, auditable, and operationally sustainable enterprise platforms, covering data processing, model training, fine-tuning, evaluation, deployment, online inference, monitoring, and lifecycle management. The position focuses on AI architecture, distributed computing, GPU resource management, Kubernetes, MLOps/LLMOps, model serving, cloud-native engineering, security hardening, reliability, and business continuity. In this job description, “hardness engineering” is interpreted as security hardening, reliability engineering, and resilience engineering for AI infrastructure . Item Details Job title AI Architecture & Infrastructure Engineer Function AI Platform, Technology Architecture, Cloud Infrastructure, and Machine Learning Engineering Reporting to Head of AI Platform / Head of Technology Architecture Location [City / Working arrangement] Employment type Full-time Number of openings [Number] Key Responsibilities 1. AI Architecture and Platform Engineering Design and deliver the architecture of the bank’s AI infrastructure and platform, covering data, training, fine-tuning, evaluation, model registry, model serving, inference gateway, identity, auditability, observability, and operations. Develop architectures for offline training, batch prediction, real-time inference, high-concurrency serving, and multi-model collaboration. Establish architecture standards, interface specifications, technical baselines, and platform evolution roadmaps. 2. Model Training and Data Processing Infrastructure Build repeatable, scalable, and auditable model-training infrastructure for supervised learning, deep learning, large language models, and other machine-learning workloads. Own training orchestration, dataset management, data preparation, feature processing, experiment tracking, model evaluation, resource scheduling, training logs, metric collection, and model-artifact management. Optimize data, compute, network, storage, and communication efficiency for training workloads. 3. Model Fine-Tuning and Large-Model Engineering Design and implement infrastructure for fine-tuning large language models and domain-specific models, supporting instruction tuning, supervised fine-tuning, parameter-efficient fine-tuning, and relevant preference-optimization workflows. Understand and apply techniques such as LoRA, QLoRA, adapters, quantization, distillation, mixed precision, gradient accumulation, checkpoint management, and distributed training. Integrate these capabilities into standardized training, evaluation, release, and rollback processes. Work with data, model-risk, business, and security teams to ensure that training data and model artifacts meet internal governance requirements. 4. Kubernetes and Cloud-Native Platform Engineering Build and operate AI workloads on Kubernetes (K8s), including containerization, Pod scheduling, GPU allocation, node-pool management, autoscaling, job queues, network policies, service discovery, storage orchestration, namespace isolation, quota management, and multi-tenant governance. Work with Helm, Operators, Ingress, service mesh, CI/CD, GitOps, Terraform, or equivalent technologies. Troubleshoot reliability, performance, scheduling, and resource-isolation issues for training and inference workloads running on Kubernetes. 5. Model Serving and Inference Optimization Build a unified model-serving platform and model gateway for conventional machine-learning models, deep-learning models, and large language models. Own model containerization, version management, canary releases, A/B testing, rate limiting, circuit breaking, caching, batch inference, asynchronous invocation, and multi-model routing. Continuously optimize inference latency, throughput, GPU/CPU utilization, memory consumption, concurrency, and cost per request. Experience with model quantization, continuous batching, KV cache, inference caching, or other acceleration techniques is preferred. 6. MLOps/LLMOps and Delivery Automation Build automated workflows covering data preparation, training, fine-tuning, evaluation, model registration, approval, release, and production monitoring. Maintain traceability among model versions, dataset versions, code versions, configuration parameters, and experiment results to ensure reproducibility, rollback, and auditability. Provide standardized SDKs, APIs, templates, pipelines, and self-service tools for data-science, machine-learning, and application teams. 7. Security Hardening, Reliability, and Resilience Engineering Establish security baselines and production-reliability mechanisms for the AI platform, including authentication, fine-grained authorization, secrets and key management, network segmentation, data masking, sensitive-data protection, image and dependency scanning, software supply-chain security, runtime protection, vulnerability remediation, audit trails, and anomaly detection. Define service-level objectives, capacity-management practices, incident-response procedures, backup and recovery plans, business-continuity controls, disaster-recovery failover, and regular resilience testing. Improve recoverability under traffic surges, dependency failures, hardware failures, model anomalies, and security incidents. 8. Observability and Production Operations Build unified observability across infrastructure, training jobs, model services, data pipelines, and business indicators. Use logs, metrics, distributed tracing, model-quality indicators, data-drift detection, performance metrics, and cost metrics to diagnose platform and model issues. Participate in production on-call rotations, major-incident reviews, capacity planning, performance testing, and platform upgrades, driving root-cause remediation rather than temporary fixes. 9. Technical Leadership and Cross-Functional Delivery Own technical design reviews, proof-of-concept validation, solution decomposition, production acceptance, and documentation for key initiatives. Mentor software, machine-learning, platform, SRE, data, security, and risk engineers. Participate in technical hiring, interviews, code reviews, and architecture reviews, helping the organization build standardized, reusable, and continuously evolving AI engineering capabilities. Minimum Qualifications Capability area Requirements Education Bachelor’s degree or above in Computer Science, Software Engineering, Artificial Intelligence, Network Engineering, Information Security, or a related field. Experience 5+ years of experience in AI platforms, cloud platforms, distributed systems, MLOps/LLMOps, platform engineering, or related infrastructure development. Experience in banking, financial services, or other highly regulated industries is preferred. Programming Strong proficiency in at least one of Python, Go, Java, or C++, with solid knowledge of data structures, algorithms, concurrency, networking, and system design. Linux and cloud-native engineering Hands-on experience with Linux, Docker, Kubernetes/K8s, container networking, service discovery, CI/CD, GitOps, and infrastructure as code. Kubernetes expertise Understanding of K8s scheduling, Deployment, StatefulSet, Job, CronJob, Service, Ingress, ConfigMap, Secret, RBAC, NetworkPolicy, resource quotas, and autoscaling. Experience with GPU workloads and multi-tenant clusters is preferred. Model training Understanding of data preparation, training orchestration, distributed training, mixed precision, checkpointing, experiment tracking, model evaluation, and model registration. Model fine-tuning Experience with supervised fine-tuning, instruction tuning, LoRA/QLoRA, adapters, quantization, distillation, parameter-efficient fine-tuning, or related large-model training techniques. AI frameworks Experience with at least one of PyTorch, TensorFlow, JAX, Hugging Face Transformers, DeepSpeed, FSDP, or equivalent frameworks and tools. Model serving Understanding of model deployment, online inference, model gateways, version control, canary releases, continuous/dynamic batching, caching, rate limiting, circuit breaking, and inference-performance optimization. Data and storage Experience with one or more of object storage, relational databases, NoSQL, message queues, data lakes, or feature stores, together with an understanding of data access, lineage, and version governance. Security and reliability Experience with IAM, secrets management, network segmentation, vulnerability management, supply-chain security, auditability, disaster recovery, incident response, and security hardening. Engineering discipline Strong focus on automation, testing, observability, documentation, code quality, change management, and production operations; able to own delivery from design through production. Communication Able to communicate effectively with business, technology, data, risk, compliance, audit, and information-security stakeholders, translating complex technical issues into clear solutions and decisions. Preferred Qualifications Experience with core banking systems, financial data platforms, risk management, anti-money laundering, customer service, intelligent operations, or other financial AI use cases. Experience with GPU clusters, NVIDIA technologies, distributed training, inference acceleration, model gateways, retrieval-augmented generation, AI agents, model security, or AI developer platforms. Familiarity with Terraform, Helm, Argo CD, Prometheus, Grafana, OpenTelemetry, Kafka, Redis, PostgreSQL, object storage, or major cloud services. Demonstrated experience building an AI platform from the ground up, scaling platform adoption, or handling major production incidents. Relevant certifications in cloud, Kubernetes, information security, data, or related technologies are a plus. Expected Outcomes and Success Measures Area Expected outcomes Platform delivery Establish a unified platform for AI training, fine-tuning, evaluation, deployment, and inference, reducing duplicated implementation. Training efficiency Improve reproducibility, resource utilization, experiment management, and delivery speed for training and fine-tuning workloads. Inference performance Optimize model-serving latency, throughput, memory utilization, concurrency, and cost per request. Production reliability Improve availability and recoverability through automation, observability, capacity management, and resilience testing. Security and governance Establish security baselines, access controls, audit records, and a closed-loop vulnerability-remediation process. Developer experience Provide standardized SDKs, templates, APIs, pipelines, and documentation to shorten the path from model development to production. What We Offer Join our banking AI infrastructure team and contribute to enterprise-grade AI platforms designed for production-scale adoption. You will collaborate with AI, cloud, data, security, risk, and business teams while working on highly reliable, secure, and performant systems. The role provides opportunities to influence critical technical decisions, platform evolution, and engineering capability development. Compensation, benefits, training, and career development are subject to the bank’s internal policies and applicable local regulations. Recruitment Process The recommended process includes resume screening, an initial technical interview, a system-design and AI-infrastructure interview, a model-training and fine-tuning interview, a Kubernetes/cloud-native technical exercise or discussion, cross-functional interviews, risk and compliance discussions, a management interview, and background checks. The bank may adjust the process based on seniority, location, and internal recruitment policies. Application Materials Please submit an English resume describing AI platforms, model-training, model-fine-tuning, Kubernetes, distributed-systems, cloud-native infrastructure, or security-hardening projects you have led or contributed to. Where possible, include project scale, technology stack, personal responsibilities, training or inference metrics, resource utilization, availability targets, cost-optimization results, and incident-management experience. Please anonymize confidential information relating to banks, customers, or internal systems. Suggested Level Naming Level Suggested title Mid-level AI Infrastructure Engineer Senior Senior AI Architecture & Infrastructure Engineer Principal AI Platform Architect / AI Infrastructure Principal Technical leadership Head of AI Infrastructure / AI Infrastructure Engineering Lead Before publishing: Complete the seniority level, compensation range, location, reporting line, on-call requirements, primary cloud environment, main models and frameworks, and applicable local compliance requirements. Document author: Manus AI References This job description was prepared based on the hiring requirements provided and does not cite external sources. Full-time

Наблюдалась 2026-09-15, впервые 2026-09-11, источник — Indeed.

Открыть у работодателя