Cohesive Technologies

Location: Remote
Experience: 7+ Years
Job Type: Full-Time / Contract
Specialization: Site Reliability Engineering, DevOps, Observability and Cloud Platforms
About the Role
Join an Observability team responsible for designing, building and operating enterprise platforms for logging, metrics, tracing and alerting across large-scale cloud infrastructure. As a Senior/Lead Site Reliability Engineer, you will play an important role in improving platform reliability, scalability and operational excellence.
This opportunity is ideal for an experienced Site Reliability Engineer, Platform Engineer or DevOps professional who enjoys working with complex infrastructure and modern observability technologies. You will help engineering teams gain better visibility into applications and infrastructure while improving how systems are monitored, diagnosed and supported.
The role offers the opportunity to work with technologies including Splunk, Elasticsearch, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Kafka and Terraform.
Key Responsibilities
Enterprise Observability Platforms
Design, deploy and operate enterprise observability platforms supporting large-scale cloud infrastructure. Build reliable solutions for collecting, processing, searching and visualizing logs, metrics and traces.
Splunk Engineering
Build and maintain Splunk Enterprise and Splunk Cloud infrastructure, including Indexers, Search Head Clusters, Heavy Forwarders and Deployment Servers.
Develop searches, dashboards and analytics using Splunk SPL. Ensure Splunk environments remain scalable, reliable and available for enterprise workloads.
Elasticsearch and Log Analytics
Deploy and operate large-scale Elasticsearch clusters for log analytics and search. Support ELK-based environments and develop effective Kibana dashboards and visualizations for infrastructure and application monitoring.
Distributed Tracing
Design, deploy and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
Build and maintain end-to-end tracing pipelines, instrumentation standards and trace retention strategies. Help engineering teams understand application dependencies and identify performance or reliability issues across distributed services.
Monitoring and Visualization
Scale and maintain Prometheus, Grafana, Kafka, Tempo and OpenTelemetry-based monitoring solutions.
Develop dashboards, alerts, analytics and trace visualizations using Splunk SPL, Grafana, Kibana and Tempo.
Automation and Infrastructure as Code
Automate infrastructure provisioning and operational processes using Terraform and configuration management tools.
Build repeatable and reliable infrastructure workflows that reduce manual effort and improve operational consistency.
Required Skills and Qualifications
Professional Experience
Candidates should have at least 7 years of experience in Site Reliability Engineering, Platform Engineering or DevOps.
The ideal candidate will have experience supporting production environments and working with enterprise-scale infrastructure where reliability, scalability and availability are critical.
Observability Experience
Strong experience with modern observability technologies is required, including:
- Splunk Enterprise or Splunk Cloud
- Splunk SPL
- Elasticsearch and ELK
- Kibana
- Prometheus
- Grafana
- Grafana Tempo
- OpenTelemetry
- Distributed tracing
- Kafka
- Metrics, logs and traces
Infrastructure Automation
Experience with Terraform and Infrastructure as Code is required. You should be comfortable automating infrastructure and operational processes rather than relying heavily on manual configuration.
Programming and Scripting
Programming or scripting experience with Python, Go, Ruby or Bash is preferred. Strong scripting ability will help you automate repetitive tasks and improve the reliability of operational workflows.
Preferred Qualifications
Splunk certification is preferred.
Experience with Kubernetes, Docker, Linux and major cloud platforms such as AWS, Azure or GCP will be an advantage.
Additional experience with Ansible, Consul, CI/CD pipelines and service mesh technologies is also desirable.
Candidates who have worked in FedRAMP or other regulated environments will be particularly valuable for enterprise environments with strict security, compliance and operational requirements.
Technology Stack
The technology environment includes:
Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible and Consul.
Why This Role Stands Out
This is more than a traditional infrastructure operations position. The role combines Site Reliability Engineering, observability, cloud infrastructure and automation.
You will have the opportunity to influence observability standards, improve monitoring capabilities and help engineering teams understand the health and performance of complex distributed systems.
For experienced SRE and DevOps professionals, this role provides an opportunity to take technical ownership and lead initiatives that directly contribute to reliability and operational maturity.
Work Model
This is a remote opportunity.
The successful candidate should be comfortable working with distributed engineering, infrastructure and platform teams. Strong communication skills and the ability to explain technical issues and recommendations clearly will be important.
Who Should Apply?
This position may be a strong fit for professionals with backgrounds in:
- Site Reliability Engineering
- Platform Engineering
- DevOps
- Cloud Infrastructure
- Observability Engineering
- Infrastructure Automation
- Production Engineering
Candidates should have strong hands-on experience with Splunk and modern observability practices. Experience with logs, metrics and traces, along with Terraform and distributed tracing technologies, will be particularly relevant.
If you are looking for additional technology opportunities, Check out other positions.
If you want to discuss your career direction, Let’s discuss your next career move.
Frequently Asked Questions
1. What does a Senior/Lead Site Reliability Engineer do?
A Senior/Lead SRE focuses on improving system reliability, availability, observability, scalability and operational efficiency across enterprise infrastructure.
2. How much experience is required?
The position requires 7 or more years of experience in Site Reliability Engineering, Platform Engineering or DevOps.
3. Is Splunk experience required?
Yes. Hands-on experience with Splunk Enterprise or Splunk Cloud and strong knowledge of Splunk SPL are important requirements.
4. What observability tools are used?
The role involves Splunk, Elasticsearch, Kibana, Prometheus, Grafana, Grafana Tempo and OpenTelemetry.
5. Is distributed tracing part of the job?
Yes. Distributed tracing is a key responsibility, with Grafana Tempo and OpenTelemetry used to build and support tracing solutions.
6. Is Terraform required?
Yes. Experience with Terraform and Infrastructure as Code is required for the role.
7. What programming languages are useful?
Experience with Python, Go, Ruby or Bash is preferred.
8. Is Kubernetes experience required?
Kubernetes is listed as a preferred qualification rather than a mandatory requirement.
9. Which cloud platforms are relevant?
Experience with AWS, Azure or GCP is preferred.
10. Is the position remote?
Yes. This is a remote opportunity.
11. What observability experience should candidates highlight?
Candidates should highlight experience working with metrics, logs and traces and building monitoring and alerting solutions at scale.
12. Is Splunk certification mandatory?
No. Splunk certification is preferred but is not listed as a mandatory requirement.
13. Can DevOps professionals apply?
Yes. Experienced DevOps professionals with strong observability, infrastructure automation and reliability engineering experience may be a good fit.
14. Is experience with regulated environments useful?
Yes. Experience supporting FedRAMP or other regulated environments is preferred.
15. Where can I find other technology jobs?
You can Check out other positions for additional technology and professional opportunities.
Ready to Apply?
If your experience matches the requirements and you are ready for your next challenge in Site Reliability Engineering and Observability, this opportunity could be a strong match.
Before applying, review your resume and make sure your experience with Splunk, Elasticsearch, Grafana, OpenTelemetry, Terraform, Kafka and cloud infrastructure is clearly presented.
To apply for this job email your details to RishiB@Cohetech.com
