What you will do:
- Infrastructure & Platform Operations: Deploy and operate platform services on Kubernetes using Helm, ArgoCD, and GitOps practices.
- Infrastructure as Code: Manage and provision our infrastructure with Terraform.
- Performance & Scale: Tune what we run, so it keeps up as the platform grows. The bigger our platform gets, the more tuning it needs.
- Self-Service Tooling & Automation: Build in-house tools that turn manual platform work into self-service platforms.
- Troubleshooting & Observability: Set up monitoring, dashboards, and alerts with Prometheus and Grafana. Debug how our systems behave under load, and write the runbooks.
- Incident Response: Handle production incidents, find the real root cause, and take part in post-mortems.
- Collaboration: Work with data engineers, data scientists, business intelligence analysts, and software engineers to unblock them and make the platform easier to use.
What you will need:
- Education: Bachelor's degree or equivalent experience in Computer Science, Information Technology, Engineering, or related fields.
- Programming: Strong programming skills (Python and/or Java preferred).
- Systems & Kubernetes: Strong Linux and networking basics. Able to deploy and manage an application on Kubernetes with Helm and ArgoCD / GitOps practices.
- DevOps & CI/CD: Experience with DevOps practices, and able to build and maintain a CI/CD pipeline.
- Observability: Able to set up monitoring and alerts with Prometheus and Grafana.
- IaC: Experience managing infrastructure with Infrastructure as Code.
- Analytical Mindset: You can explain your troubleshooting steps and justify your technical choices, not just show us the fix.
- Language & Location: Proficient in English and Thai, and based in or close to Thailand, or willing to relocate.
It'd be Great if you have:
- Hands-on experience with a data stack like ours — Apache Airflow, lakehouse architectures, the Hadoop ecosystem, Databricks, or a data warehouse. You don't need to be a deep expert, but if you have worked on this kind of stack before, you will find your way around much faster.
- Understanding of how a platform component behaves at large scale — for example, what to tune when Airflow runs thousands of DAGs.
- Understanding of how Spark or Flink runs a job, and how you would process 100 GB+ of data efficiently.
- Experience with an OLAP data warehouse, e.g. ClickHouse, Doris, StarRocks, Redshift, or SQL Server.
- Experience building an in-house service yourself, front-end and backend.
- Experience with Apache Iceberg or another open table format.
- Experience with managing infrastructure for BI Tools (e.g. Tableau, PowerBI, Redash, Metabase, Superset, etc.)
