Deputy IT Infrastructure Lead / Infrastructure Engineer - Saigon Precision
(2026-03)
- Designed, implemented, and optimized enterprise monitoring systems using Zabbix and Grafana deployed via Docker Compose and Ansible
- Developed custom monitoring dashboards in Grafana without using prebuilt templates, including customized visualization for infrastructure, servers, and network devices.
- Implemented real-time network topology and bandwidth monitoring using Grafana Network Weathermap plugin to visualize traffic flow and connectivity between network devices.
- Customized Zabbix metrics, monitoring parameters, triggers, and alert thresholds to improve infrastructure observability, monitoring accuracy, and incident detection efficiency.
- Integrated monitoring and alerting systems with Microsoft Teams using Teams Workflows and webhook-based notification pipelines.
- Managed and maintained firewall systems, VMware virtualization infrastructure, Active Directory, VPN systems, Microsoft Intune (basic administration), and internal IT infrastructure services.
- Manage virtual machines and migrate them from VMWare on-premise to a CMC cloud environment.
- Managed backup and disaster recovery operations using Veeam Backup & Replication, including offsite backup strategies with S3-compatible object storage.
- Monitored backup integrity, retention policies, and recovery validation procedures to ensure infrastructure reliability and business continuity.
- Administered VMware virtual machines running Oracle Database and Microsoft SQL Server workloads, including monitoring, maintenance, and operational support.
- Managed IT asset inventory including PCs, network devices, and server room equipment.
- Coordinated daily operations and technical support activities for the helpdesk team.
- Proposed and implemented operational improvements, technical documentation, and team upskilling initiatives to improve troubleshooting capability and infrastructure management efficiency.
Systems Deployment Engineer - Thi Thien Solutions Technology Corporation
(2023-11 - 2026-03)
- Deploying SQL Server Failover Cluster Instance (FCI): Configured high availability for SQL Server 2022 on Windows Server 2022 across 2-node clusters.
- Storage Management (HPE MSA 2060): Implemented iSCSI Direct Attach, MPIO, and Storage Tiering (SSD/HDD) to optimize performance and redundancy.
- Disaster Recovery: Successfully restored critical database services by reconfiguring SAN volumes and Cluster nodes after a major storage hardware failure.
- Recovered ESXi 8 root access for Post Office General Hospital, restoring critical hospital systems after 10 days of outage within 2 hours, preventing further operational disruption.
- Recovered ESXi 7 password for Vietsovpetro Resort Ho Tram, helping the resort's system to resume operation (processing time (Downtime less than 1 hour).
- Resolved HA synchronization issues on a production HPE SimpliVity cluster for Saigon Giai Phong Newspaper.
- Troubleshoot and replaced server, SAN, and tape library hardware, ensuring minimal service downtime.
- Install, configure, test, and troubleshoot customer systems including VMware, Proxmox, Linux, Windows Server, and domain controllers.
- Install and configure Servers, SANs, and Tape Libraries.
- Performing troubleshooting tasks such as password reset and IP reset for IBM flash systems (v5000, v7000, v3700).
- Diagnosed and resolved complex hardware/software issues for SAN storage systems (Dell, IBM, 3PAR, MSA) ensuring 99.9% system uptime.
- Troubleshot and rebuilt RAID arrays following disk failures, ensuring data integrity and system recovery.
- Handling firewall, switch, and network issues for customers, including server
- Supporting Hung Vuong Hospital in installing and configuring VMWare ESXi 8. Installing and configuring Oracle Linux and Oracle database.
- Installed CUDA drivers and configured TensorFlow, PyTorch, and Jupyter Notebook with GPU support for AI/ML workloads.
- Support rebuilding RAID for customers
Freelance System Engineer - Freelance
(2025-04 - 2025-04)
Short-term engagement (April 2025) and short-term engagement (October 2025):
- Troubleshot Kubernetes cluster issues preventing worker nodes from joining the cluster.
- Restored production workloads by redeploying existing Kubernetes manifests and configurations.
- Troubleshot OpenStack deployment issues during installation using Kolla-Ansible.
- Investigated service startup failures, configuration errors, and dependency issues during deployment.
- Assisted in redeploying existing OpenStack configurations to successfully complete the deployment and bring services online.
IT Systems Specialist - Delta Engineering Corporation (Now is DEC Engineering Joint Stock Company)
(2020-09 - 2022-04)
- Deploy, upgrade, and configure network devices (WiFi, cameras, switches, routers, NAS, servers, thin clients).
- Implement a system for monitoring devices in a network using Zabbix, Grafana
- Implementing a CEPH distributed storage system for use with OpenStack and Proxmox.
- Deploy and operate the OpenStack system.
- Deployment and administration of Proxmox cluster systems.
- Installed Apache Guacamole Remote gateway.
- Deploy and operate a domain controller system.
- Build and operate a private cloud with Nextcloud and integrate SSO with ADFS.
- Integrated ADFS for Microsoft Office 365.
- Implement VDI using Windows Server for centralized storage with a thin client device for end users.
- Deploying and operating VMWare ESXi, Vcenter
- Configure firewall devices (Sophos, Fortinet, pfSense)
- Update the system with the latest patches.
- implemented High Availability (HA) architectures for MariaDB and PostgreSQL databases using HAProxy, Galera Cluster, and Patroni to ensure fault tolerance, load balancing, and automatic failover.
- Maintained computers, network devices, cameras and servers
- Server management, virtual machines, and data backup and restore
- Automated routine system administration tasks using Bash and Python, reducing manual effort by 80%.
- Write deployment documentation for administrators and user guides.
- Supported resolved network-related issues.
- Perform other tasks assigned by superiors.
System Operation - Berjaya Gia Thinh Investment Technology Joint Stock Company
(2019-01 - 2019-12)
System operation for Vietlott - Main responsibilities:
- Monitoring and checking the server and service systems.
- Perform daily and weekly backups to tape.
- Sent reports to relevant departments.
- Relocated tapes to PDC and DRC.
- Managed start and close systems for Vietlott.
- Installed mail server on CentOS 6.
- Operated OSS (Online Sell Server) on OpenVMS.
- Performing system operation tasks during lottery draws.
Operator (NOC) Trainee - ONLINE MOBILE SERVICES JOINT STOCK COMPANY (MoMo)
(2018-01 - 2018-08)
Main responsibilities
- Monitored servers and services (Nagios, Grafana, Web).
- Supported various departments (Customer Care, Product Development, Distribution System, Promotion).
- Performing tasks such as logging errors, handling user-reported or system-recorded issues.
- Monitors the company's network, keeps systems and services running and has no problems.
- Experience the work environment of the business