22/09/2026
CNTT - Phần cứng / Mạng, Bưu chính viễn thông, Điện / Điện tử / Điện lạnh / Điện công nghiệp
Nhân viên chính thức
Cạnh tranh
Trên 3 Năm
Nhân viên
20/10/2026
- Evaluate and implement the platform: Compare shortlisted open-source and commercial solutions through practical pilots, document capabilities/integration effort/operating costs/limitations, recommend a stack to the Tech Lead, and implement the approved solution.
- Build the hosting environment: Configure servers, virtual machines, operating systems, databases, network connectivity, certificates, storage, and supporting services; maintain separate test and production environments with repeatable deployment procedures.
- Implement infrastructure monitoring: Deploy central monitoring and customer-site collectors; configure discovery, monitoring templates, dashboards, thresholds, dependencies, maintenance windows, and checks for networks, Windows/Linux servers, virtualization, and agreed cloud services.
- Integrate monitoring with service operations: Build or configure API and webhook integrations so alerts create/update correct tickets and notify appropriate engineers; implement duplicate-event handling, retries, recovery updates, and notification continuity during ticketing outages.
- Implement customer separation and secure access: Configure customer-scoped permissions, SSO/MFA, service accounts, secrets handling, and audit records.
- Automate repeatable work: Develop reviewed scripts and Ansible workflows for installation, configuration, inventory, diagnostics, and approved operational tasks; use version control, scoped credentials, execution records, and rollback procedures.
- Validate production readiness: Test backups and restoration, supported high availability, collector disconnections, platform failures, capacity, and customer onboarding/offboarding; track defects and provide evidence against agreed acceptance criteria.
- Prepare operational handover: Deliver deployment documentation, diagrams, runbooks, troubleshooting guides, and onboarding checklists; train operations staff, mentor the System Intern, and provide technical escalation after launch under an agreed support arrangement.
- At least 3 years of relevant systems or infrastructure engineering experience, including hands-on implementation of production environments.
- Experience delivering at least one monitoring, infrastructure management, or service-management platform through deployment and operational handover.
- Strong Linux administration and practical Windows Server knowledge.
- Working knowledge of TCP/IP, DNS, routing, firewalls, TLS, virtualization, and cloud infrastructure.
- Practical experience with monitoring configuration, alert tuning, and incident workflows.
- Ability to integrate systems using REST APIs, webhooks, and scripting in Python, PowerShell, or Bash.
- Experience with Git and configuration automation, preferably Ansible.
- Ability to document designs, explain technical tradeoffs, and coordinate implementation with colleagues, customers, and vendors.
- Ability to read and work with English technical documentation.
Preferred experience
- Zabbix or ManageEngine OpManager; iTop, ServiceDesk Plus MSP, or comparable CMDB/ITSM platforms.
- NetBox, paging platforms, secrets management, or privileged-access tooling.
- VMware, Hyper-V, or another virtualization platform; AWS, Azure, or GCP.
- Database administration, container deployment, or infrastructure as code.
- Experience in an MSP, managed hosting, or multi-customer environment.