We are seeking an Observability DevOps Engineer, to join our Automation team and help us scale our next-generation network monitoring and telemetry platform. In this role, you will design, build, and optimize high-volume data pipelines that collect telemetry from thousands of network devices across diverse technologies (APIs, SNMP, streaming telemetry) and transform this data into actionable insights for our organization.
You'll work with our modern observability stack—Telegraf → Kafka → ClickHouse → Grafana— managing everything from data ingestion and streaming buffers to distributed database clusters and real-time visualization dashboards. This is a hands-on technical role where you'll shape the platform roadmap, drive performance optimization at scale, and empower teams with reliable, self-service monitoring capabilities.
Key Responsibilities
• Design and maintain scalable telemetry data pipelines processing high-volume network metrics, logs, and events across our Telegraf → Kafka → ClickHouse architecture
• Manage and optimize our 3-node ClickHouse cluster, including complex query optimization, partitioning strategies, materialized views, and distributed table design
• Develop and maintain data collection workflows using Telegraf, integrating with APIs, SNMP, and various network/servers device protocols
• Build and enhance Grafana dashboards and alerting mechanisms to provide real-time visibility into infrastructure performance
• Automate infrastructure and platform operations using IaC principles, configuration management, and CI/CD pipelines
• Collaborate with operations teams to understand monitoring requirements and translate them into scalable technical solutions
• Troubleshoot and resolve platform incidents, performing root-cause analysis and implementing preventive measures
• Document platform architecture, runbooks, and operational procedures to enable team scalability
Required Skills
Core Technical Competencies
• Programming & Scripting: Strong proficiency in Python and Go for building platform tools, automation, and data processing
• SQL & Database Expertise: Advanced SQL skills with experience writing and optimizing complex queries for analytical databases
• API Development & Integration: Experience designing, developing, and consuming REST APIs for data collection and platform services
• Data Structures: Deep understanding of JSON, XML, and other data formats used in telemetry and monitoring
• Linux System Administration: Solid Linux administration skills for managing production infrastructure
• Version Control: Proficiency with Git/Bitbucket for collaborative development and infrastructure-as-code workflows
• Data Pipeline Architecture: Experience designing and operating high-throughput data pipelines and stream processing systems
Collaboration & Process Tools
• Jira and Confluence for project management and documentation
• ServiceNow for ticket management and operational workflows
Nice-to-Have Skills
• Apache Kafka: Experience with Kafka for building streaming data platforms and managing topics, partitions, and consumer groups
• ClickHouse: Hands-on experience with ClickHouse or similar columnar analytical databases (TimescaleDB, Druid)
• Grafana & Loki: Experience building dashboards, alerts, and log aggregation solutions
• OpenTelemetry (OTEL): Knowledge of modern observability standards and instrumentation
• Infrastructure as Code: Experience with Terraform, Ansible, or similar tools for automating infrastructure provisioning
• Containerization: Basic knowledge of Docker and container orchestration concepts
• Shell Scripting: Proficiency in Bash and PowerShell for automation tasks
• Prometheus: Familiarity with Prometheus metrics and the observability ecosystem
• SRE Principles: Knowledge of SLOs, SLIs, error budgets, and reliability engineering practices
• Effective Communication & Vendor Management: Ability to handle escalations, coordinate incident resolution, and maintain clear communication with third-party monitoring service providers.
What Makes This Role Exciting
• Work with cutting-edge technology: Build and scale a modern observability platform using Kafka, ClickHouse, and Grafana—technologies at the forefront of high- performance data engineering
• Massive scale and impact: Process telemetry from thousands of network devices and provide insights that directly impact Infrastructure reliability and performance across
the organization
• Platform thinking: Design systems that enable other teams to self-serve their monitoring needs, creating a true platform-as-a-service experience
• Technical ownership: Drive architectural decisions, optimize performance, and shape the future direction of our observability infrastructure
• Diverse technology stack: Work across the full telemetry pipeline—from device integration to distributed storage to real-time visualization
Hybrid Scheme in Munro