Jump to content

Senior/Lead Site Reliability Engineer Observability - San Jose, CA | TCS Job


 Share

Job Opportunity Details

Type

Full Time

Salary

Not Telling

Work from home

No

Weekly Working Hours

Not Telling

Positions

Not Telling

Working Location

San Jose, CA, San Jose, CA, United States   [ View map ]
Must Have Technical/Functional Skills:
• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
• Hands-on experience administering Splunk Enterprise or Splunk Cloud.
• Strong knowledge of Splunk SPL.
• Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.
• Experience implementing metrics, logs, and traces as part of a modern observability strategy.
• Experience with Terraform and Infrastructure as Code.
• Programming experience in Python, Go, Ruby, or Bash. 
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.
• Experience supporting FedRAMP or regulated environments. 
Technology Stack
Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.
 
Roles & Responsibilities:
• Design, deploy, and operate enterprise observability platforms.
• Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers.
• Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
• Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.
• Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
• Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.
• Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
• Automate infrastructure using Terraform and configuration management tools. 

Nice to have skills:
• Splunk certification.
• Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
• Experience supporting FedRAMP or regulated environments.

In order to comply with U.S. laws and regulations applicable to this position, the person(s) hired must possess the ability to obtain US Security Clearance which requires that the person be a U.S. Citizen, a U.S. Permanent Resident (i.e., a “Green Card Holder”), or a Political Asylee or Refugee.

Salary Range: $64,000 - $130,000 a year
#LI-CM2


More Information

Application Details

  • Organization Details
    TCS / Tata Consultancy Services
 Share


User Feedback

Recommended Comments

There are no comments to display.

Join the conversation

You are posting as a guest. If you have an account, sign in now to post with your account.
Note: Your post will require moderator approval before it will be visible.

Guest
Add a comment...

×  Pasted as rich text.   Paste as plain text instead

  Only 75 emoji are allowed.

×  Your link has been automatically embedded.   Display as a link instead

×  Your previous content has been restored.   Clear editor

×  You cannot paste images directly. Upload or insert images from URL.

Loading...
×
×
  • Create New...