Devops
- Experience: 1 - 3 Years
- Location: Indore
Requisition Description
He acts as the first line of support for environment issues within the BDA team and is accountable for monitoring cluster performance and health.He operates as a hybrid DevOps + Dev Support engineer with growing responsibilities, helping the team move toward improved reliability, automation, and cluster stability.He is also responsible for creating runbooks for deployment guidelines and documenting issues that have been troubleshooted.
Roles and Responsibilities Cluster Deployment & ManagementDeploy fresh BDA clusters end-to-end using third-party components and CT Components.Maintain cluster lifecycle operations (upgrade, scale-up/scale-down, patching using existing DevOps or Production tools (able to develop scripts where required).Ensure cluster configurations follow production standards and best practices (required in dev internally). Maintain environment parity between Dev and dev staging (In case of issues, limited responsibility, but should not allow tweaking's without alignments hence enforce guardrails across) clusters.Monitoring, Alerting & Health ChecksContinuously monitor all dev clusters for performance, availability, and resource consumption.Set up and maintain proactive alerts (CPU, RAM, disk, JVM, pods, solr indexing).Investigate alerts and take corrective actions before they impact development.Maintain health dashboards and generate weekly cluster health reports to stakeholders.Troubleshooting & Issue ResolutionAct as the first-level point of contact for all Dev environment issues. Should also be able to troubleshoot QA and L3 issues (future expectation).Troubleshoot cluster failures, deployment issues, indexing problems, ingestion errors, and service outages.Perform initial triage and provide detailed, actionable analysis.Observability & Logging ImprovementsProvide suggestions for observability and logging improvements to Dev and HM teams.Ensure logs, metrics, alerts, and dashboards are up-to-date and accurate.Monitor logs and dashboards for cluster health and component-level metrics.Logs should be clean and in case of error report same to stakeholders.Environment Readiness for DevelopersEnsure all clusters are stable, updated (New Gathr versions in alignment with stakeholders), and ready for releases and validation cycles.Support developers with environment issues, configuration needs, and component access.Validate and certify environments before major releases.Documentation & RunbooksDocument deployment procedures, troubleshooting steps, runbooks, knowledge base, and environment setup guides and publish it to Shekhar & team.Maintain a clear configuration inventory and version matrix for each cluster.Create runbooks for repetitive operational tasks and L3 incident handling. Ultimate goal is to have him point of contact for Production L3 support. Own end-to-end problem resolution for production-level issues. (Future)Provide deployment guidelines and support to dependent teams.Performance Optimization & Capacity ManagementAnalyze cluster performance and identify bottlenecks.Work with team to fine-tune components (Solr, Doris, Milvus, Delta, Kafka, etc.).