Working in the Hardware Team at OCF is a role that is both challenging and incredibly rewarding.
We are responsible for building and assisting in troubleshooting the High Performance Computing (HPC) systems that our customers rely on.
In this blog post, I want to give you an insight into what it’s like to be a part of the team, the kind of work we do, and the skills it takes. Whether you’re interested in the hardware side of HPC, considering a similar role, or just curious about what goes into delivering cutting-edge tech, I hope this gives you a useful glimpse into life at OCF.
No two projects are ever quite the same in my role, but here’s a look at the typical processes that go into a project from a hardware perspective.
All projects begin with the in-house work, which consists of preparing for the site visit by creating a complete point-to-point and rack diagram based on the customer’s requirements. These documents detail how the cluster will be connected and the layout of the rack(s).
While creating the point-to-point and rack diagram, many different factors need to be considered. For example, the type of connections that each different server may have, as different servers have different priorities and require different speeds/types of connections. Some servers may only require management connections and lower speed ethernet connections; however, other servers may require high-speed InfiniBand / Ethernet and even dual connections for redundancy to reduce any risk of outages and downtime if a cable fails.
The lengths of the cables used also need to be considered to ensure no unnecessary stress is put on the cables whilst still leaving as little slack as possible to reduce rack clutter. Many things can alter the required length for a connection, starting with the overall distance between the server and switch, whether they will cable from the back or the front, if there are gaps between racks, and the distance to cable trays above/below the rack.
Once this documentation has been created, if a cluster requires benchmarks or configuration changes before heading to the site, we will assist the compute team in setting up the cluster to perform these tests or changes.
Once this in-house preparation work has been completed, we then work on-site and begin the installation. We begin by checking over the kit that has been delivered to the site to ensure that we have all the servers, switches, and cables to complete the installation.
We will then move on to checking the internal rack rails are correctly set and adjusted to the needs of deeper pieces of equipment, such as GPU nodes or storage shelves. Once these rails are set correctly, we will move on to unboxing and racking the nodes following the rack diagram created earlier.
After all the kit is racked, we are ready to move on to the cabling. Following along from the point-to-point, we typically begin with the power cables, as they are the main bulk and take up the most space within the rack. Whilst cabling, it is important to ensure that we choose the right path and secure the cables well. This ensures the customer is left with a neat rack, as well as improving airflow and making it easier to service a node if any issues arise in the future. We pride ourselves on our cabling and ensure that it is always of a high standard for our customers.
Next, we move on to any high-speed networking, as they may require extra space due to bend radii to reduce any stress on cables. When the high-speed networking and power are cabled, we then move onto 25GbE or 10GbE cables and 1GbE cables, which don’t take up lots of space and are generally quite malleable.
Finally, we install any ISLs (Inter-switch links) or uplinks to the customer network. The cabling is now complete, and we are ready to perform green light/power-on tests on all of the nodes and switches to ensure there are no faulty cables or power issues. If any issues arise during this stage, we will attempt to troubleshoot on-site, and if not, we will escalate to find a solution down the line.
After all servers have booted successfully, we then print labels for both sides of every cable run, noting the U numbers and port numbers of each end. This assists with future maintenance or alterations of the cluster.
Overall, being part of the Hardware Team at OCF is about more than just installing equipment. It’s about delivering reliable, high-performance systems that our customers can trust. Every project presents new challenges and learning opportunities, and that’s what makes the role so rewarding.
Seeing a cluster go from detailed design to a fully operational system is something I’m proud to be a part of every time.