Site Reliability Engineer - Data Engineering

Talenza · Sydney NSW 2000 · Full time
Posted today

Site Reliability Engineer - Data Engineering

KEY POINTS WE FOUND
  • Join a global trading firm's Data Engineering team as a Site Reliability Engineer.
  • Operate and improve large-scale data platforms including Kafka, HDFS, and Dremio.
  • Troubleshoot production issues and build automation for deployments.

Site Reliability Engineer - Data Platform

Sydney | Global Trading Firm

Join a global trading firm's high-performing Data Engineering team and help operate the large-scale platform that underpins trader research, simulation, reporting and decision-making.

This is a proper infrastructure role-not one for someone who has only built data pipelines on top of managed services. You'll get hands-on with the underlying platform: Kafka, HDFS, Dremio, Linux and in-house data tooling across a multi-petabyte environment processing around two million queries each day.

You'll join a small, experienced Sydney team with international engineering counterparts, giving you genuine ownership, strong mentoring and exposure to complex distributed systems at meaningful scale.

The role

  • Run, monitor and improve large-scale data platforms including Kafka, HDFS, Dremio and internally built pipelines.

  • Troubleshoot real production issues across Linux, storage, networking and distributed infrastructure.

  • Build automation and CI/CD capability to make deployments faster, safer and more repeatable.

  • Support upgrades, capacity planning, incident response and long-term reliability improvements.

  • Work closely with systems and network engineers, developers, researchers and end users to solve complex data-platform problems.

  • Help evaluate and introduce new technology as the environment continues to evolve.

What we're looking for

  • Around 2-3 years' experience in an SRE, platform, infrastructure, systems or production engineering role.

  • Strong Linux troubleshooting skills across processes, filesystems, networking, disk and memory pressure.

  • Hands-on operator experience with at least one of Kafka, HDFS or Kubernetes-you have deployed, configured, upgraded, tuned or supported the platform itself.

  • Experience operating self-managed infrastructure, whether bare metal, datacentre, self-run VMs or self-managed Kubernetes.

  • Python experience for systems automation, operational tooling, health checks or deployment workflows.

  • Exposure to Docker, Kubernetes, Helm and infrastructure-focused CI/CD.

  • A curious, pragmatic mindset and a genuine interest in understanding how complex systems behave under pressure.

Experience with Dremio, Presto, Airflow, Prefect, Ansible, Puppet, Terraform or cloud platforms would be beneficial, but it is the operational mindset and underlying Linux/infrastructure depth that matter most.

This is a standout opportunity for an engineer who wants to move beyond managed services, get close to the underlying technology and build a career operating high-scale, business-critical systems.

Stay Safe While Job Hunting

We vet all employer accounts and do our best to keep job ads safe, but scams can still occur. Be cautious when sharing personal information — never provide financial details or make payments during the application process. For extra security, use the Apply button on our site when proceeding.

Skills

0 of 23 matched
AirflowApache hadoop hdfsAutomationCi/cdCloud platforms (beneficial)Complex systems understandingDistributed systemsDockerDremioHelmIncident responseKafkaKubernetesLinuxMonitoringNetworkingOperational mindsetPrestoPuppetPythonSwitch capacityTerraformTroubleshooting

Talenza

Talenza Logo

More details

Expiring date
Site Reliability Engineer - Data Engineering | Talenza | HIA