Test & Failure Analysis (FA) Engineer
Build the Future of Artificial Intelligence in Texas
Location: Texas, USA (On-site)
Employment Type: Full-Time
Build Your Engineering Career in the U.S. AI Industry
Join one of the world's leading AI server manufacturing organizations and help build the infrastructure powering Artificial Intelligence, Cloud Computing, High Performance Computing (HPC), and Hyperscale Data Centers.
We are seeking talented Test & Failure Analysis Engineers who are passionate about solving complex hardware challenges, performing advanced diagnostics, and driving continuous improvement in AI server manufacturing.
If you thrive on technical problem solving and want to work with cutting-edge server technologies, this is your opportunity to make an impact while building a long-term engineering career in the United States.
Why You'll Love This Opportunity
- Work with next-generation AI and GPU server platforms.
- Troubleshoot some of the world's most advanced enterprise server technologies.
- Support products used by leading hyperscale cloud providers.
- Join one of the fastest-growing industries in the world.
- Collaborate with experienced engineering professionals.
- Excellent long-term career growth opportunities.
- Be part of a world-class manufacturing organization in Texas.
Position Summary
The Test & Failure Analysis Engineer is responsible for diagnosing complex hardware failures, performing Root Cause Analysis (RCA), validating AI server platforms, and supporting manufacturing operations to ensure world-class product quality and reliability.
This position works closely with Manufacturing, Test, Quality, Product Engineering, and New Product Introduction (NPI) teams to improve manufacturing yield, product reliability, and customer satisfaction.
Key Responsibilities
- Troubleshoot complex AI and Hyperscale server hardware failures.
- Perform Root Cause Analysis (RCA) using structured problem-solving methodologies.
- Analyze manufacturing failure logs and identify hardware-related issues.
- Execute manufacturing test procedures and Standard Operating Procedures (SOPs).
- Perform component isolation and hardware verification.
- Validate server functionality during production and system integration.
- Troubleshoot issues related to CPUs, GPUs, memory, storage, networking, and power systems.
- Utilize Linux command line tools for diagnostics and validation.
- Support BIOS, UEFI, BMC, and IPMI configuration and troubleshooting.
- Develop and execute corrective and preventive actions to improve manufacturing quality.
- Document failure analysis findings and engineering recommendations.
- Support New Product Introduction (NPI) and production ramp activities.
- Collaborate with cross-functional engineering teams to improve product reliability and manufacturing efficiency.
Required Qualifications
- Bachelor's Degree in Electrical Engineering, Computer Engineering, Computer Science, Mechatronics, or a related engineering discipline.
- Experience troubleshooting enterprise server hardware.
- Strong Root Cause Analysis (RCA) skills.
- Solid understanding of server architecture fundamentals.
- Linux command line proficiency.
- Experience with:
- BMC
- IPMI
- BIOS
- UEFI
- Manufacturing test process experience.
- Failure log analysis.
- Component isolation and hardware verification.
- Basic Ethernet networking knowledge.
- Experience executing Standard Operating Procedures (SOPs).
- Strong analytical and problem-solving skills.
- Ability to work in a fast-paced manufacturing environment.
Preferred Qualifications
Experience in one or more of the following is highly desirable:
AI Infrastructure
- AI Server Manufacturing
- AI Server Validation
- Hyperscale Server Platforms
- Enterprise Server Architecture
- Data Center Infrastructure
- High Performance Computing (HPC)
Advanced Hardware
- GPU Platforms
- NVIDIA HGX
- NVIDIA DGX
- AMD Instinct
- Rack Integration
- Rack-Level Validation
Engineering & Manufacturing
- Failure Analysis
- Manufacturing Test Engineering
- Manufacturing Execution Systems (MES)
- Functional Testing
- Burn-In Testing
- New Product Introduction (NPI)
- High-Volume Electronics Manufacturing
Software & Automation
- Python Scripting
- SQL
- Jira
- OpenBMC
- Redfish
Advanced Technologies (Preferred)
- PCIe
- NVLink
- DDR5 Memory
- NVMe Storage
- Liquid Cooling Systems
- OCP Platforms
Working Conditions:
- Salary Range: $60,000 – $80,000 MXN per month.
- Perks / Benefits: Accommodation and vehicle included during the assignment.
Value Proposition (Benefits)
- Visa Sponsorship: Comprehensive management and fully covered sponsorship for a work visa.
- Benefits Package: Comprehensive benefits package and medical coverage compliant with U.S. labor laws.
- Career Growth: Global professional advancement in highly sophisticated, technological industrial environments.
- Job Security: Guaranteed job security backed by international corporate support.
Sueldo: $60,000.00 - $80,000.00 al mes
Lugar de trabajo: Empleo presencial