Cloud Infrastructure & Platform Operations: Design, deploy, and maintain cloud infrastructure that meets the organization's performance, security, and cost requirements
CI/CD Automation: Build and operate continuous integration and continuous deployment of pipelines, ensuring software quality and release velocity
Monitoring & Incident Management: Maintain comprehensive observability, detect and resolve incidents quickly, and ensure high availability
Database & Data Management: Administer, optimize, and secure SQL databases; support backup, recovery, and data migration
Security & Compliance: Apply DevSecOps practices throughout the software lifecycle; ensure systems comply with security policies and applicable regulations
Continuous Improvement & Technical Mentoring: Foster a DevOps culture, continuously improve operational processes, and provide technical guidance to team members
Main Responsibilities
1. Main Responsibilities:
Design and deploy cloud infrastructure using Infrastructure as Code (Terraform, Ansible) on AWS/GCP, ensuring high availability (HA) and disaster recovery (DR) capability
Build, operate, and optimize CI/CD pipelines (GitLab CI/Jenkins/GitHub Actions) to automate the entire build, test, and deployment workflow
Deploy and manage Kubernetes clusters in production: configure workloads, autoscaling, resource quotas, and network policies
Set up comprehensive monitoring (Prometheus, Grafana, ELK Stack) for infrastructure and applications; build runbooks and on-call procedures
Administer and optimize SQL databases (PostgreSQL, MySQL) in production: automated backups, performance tuning, and safe data migration
Manage system capacity: track resource usage trends, plan scaling, and control cloud costs
Apply DevSecOps: integrate security checks into the CI/CD pipeline and GitOps workflow, and manage secrets (Vault/K8s Secrets), certificates, and access according to the principle of least privilege
Support root cause analysis of production incidents; propose and implement improvements to increase system stability
Collaborate with the software development team to design application architecture aligned with infrastructure and operational requirements
Standardize operational processes, document infrastructure, and share technical knowledge with the team
2. Other Responsiblities:
Champion DevOps, SRE, and Software Security principles across the organization, serving as a technical authority to instill a culture of excellence
Collaborate with DTI development teams to align project delivery
Partner with core IT functions, including Infrastructure, Security, and Service Management
Job Requirement
1. Educational level
Bachelor’s degree in computer science, Information Technology, or a related field
2. Professional Experience
Over 5 years of professional experience in DevOps, System Administration, or Software Engineering
Deep expertise in DevSecOps and SRE operational models. A proven track record of integrating robust security standards throughout the Software Development Lifecycle (SDLC)
Extensive expertise in microservices architecture, RESTful APIs, and the deployment of distributed systems
Direct experience in administering and operating Cloud platforms (GCP, AWS) and Containerization, Orchestration ecosystems (Kubernetes, GKE, RKE)
Proven experience in establishing and managing a DevOps Service model to support multiple concurrent product development teams
Expertise in IaC (Terraform, Ansible) and the GitOps framework (ArgoCD)
Designing and deploying comprehensive Monitoring/Observability stacks (Prometheus, Grafana, ELK/EFK, APM) for multi-cluster Kubernetes environments
Practical experience using AI coding tools in daily work
Basic understanding of prompt engineering to effectively leverage LLMs for technical tasks
AI governance mindset: clear awareness of security/compliance risks when using AI in a financial environment
3. Technical skills
Ability to work independently, take initiative, and contribute to new ideas required in a diverse, fast-paced, deadline-driven team environment.
Ability to conduct and direct research into issues and products as required.
Highly self-motivated and directed.
Keep attention to detail.
Proven analytical, evaluative, and problem-solving abilities.
Able to prioritize and execute tasks in a high-pressure environment.
Experience working in a team-oriented, collaborative environment.