面議(經常性薪資達4萬元或以上) 台北市內湖區 3年工作經驗 2天前更新
Job Purpose : Visionbay.ai is one of the leaders of AI Data Center and AI Factory in Taiwan. A System Specialist in Visionbay Architecture Team helps to design, implement and deliver CPU/GPU servers and storage that are crucial in an HPC environment.
Responsibilities
technical planning for AIDC projects
collect, compile, and synchronize server, systems, and storage related requirement
define, agree on the acceptance criteria with customers / users
provide high-level systems and storage solutions based on the requirement
perform detailed systems and storage design based on the solution
provide implementation plan or migration plan based on the above
technical delivery for new AIDC projects
install and configure the servers and storage to bring up compute resources
install and configure server-side middleware and applications
(optionally) perform configuration change over production environment based on the migration plan
perform testing, validation, and fine-tuning for the systems and storage
(optionally) on-call during change period for production environment
documentation handling until handover
generate and update the requirement, solutions, detailed design, and acceptance documents
generate and update the implementation plan, migration plan
acceptance & handover
accept the deliverables from suppliers based on the agreed criteria
deliver training sessions to internal teams
handover the infrastructure to operation teams
Competence
must have (at least 80%)
basic knowledge throughout TCP/IP model and OSI model
feel comfortable working via both CLI and GUI
able to install and configure servers and storage from scratch
RHCA or equivalent hands-on experience
expert level of knowledge and experience on operating systems such like Linux, *nix and Windows Server
expert level of knowledge and experience on hardware/software RAID, block storage
expert level of change management for production environment
expert level of knowledge and experience on server virtualization technologies including hypervisors and containers
familiar with container orchestration platform such like Kubernetes
familiar with PXE-based system administration
familiar with server hardware, firmware, BIOS, and IPMI/BMC management interfaces
familiar with public cloud services such like AWS, GCP, Azure, Alicloud
familiar with modern identity systems such like Active Directory, Entra ID, Okta
familiar with authentication and access controls such like SSO, MFA, Conditional Access, Device Trust, and Zero Trust principles
familiar with file storage, object storage
familiar with shell scripting or Python or other scripting language for job automation
familiar with web services, HTTP headers, SSL/TLS certificates
展開