Twenty-four hours of Live Optics is not a workload profile
Second in the series. The collection window is the single input that most often invalidates a sizing, and it is the one nobody negotiates.
Arrays, protocols, latency, and honest numbers.
Second in the series. The collection window is the single input that most often invalidates a sizing, and it is the one nobody negotiates.
Third in the series. IOPS, block size, read/write mix and the compression estimate — what each one is actually measuring, and where the sizing goes wrong.
It is not a faster NAS. The architectural difference is that clients talk to many storage servers at once, and almost every operational surprise follows from that.
Every parallel file system is sold on GB/s. Almost every one that disappoints in production is disappointing because of operations per second on small files.
Two numbers control whether your parallel file system behaves like dozens of servers or like one. Most sites never change them.
The performance is not the issue. The issues are upgrades, client evictions, one slow target poisoning everything, and the fact that it assumes you employ someone who knows it.
GPFS by its older name. Technically excellent, operationally more forgiving than Lustre, and sold in a way that requires you to read carefully.
There is a real gap in the middle of this market, and the thing that fills it is usually the one nobody shortlisted.
Flash removed the seek penalty that shaped twenty years of design. It did not remove the metadata bottleneck, the network limit, or the need to think about layout.
Capacity is the easy number and the wrong one to start from. Here is the order that produces a design somebody can defend.
Every shared file system struggles with millions of tiny files. Training pipelines produce exactly that, and the fix is in the pipeline rather than the storage.
They rarely stop. They degrade, partially and confusingly, in ways that present to users as 'the storage is slow' with no pattern.
Walking a namespace of a billion files with a conventional backup agent will not finish. The strategies that work look nothing like enterprise backup.
The default answer for an enterprise AI cluster should be NAS, and the burden of proof should sit with the parallel file system. Usually it is the other way round.
Not everything needs to sit on the fast tier. The interesting design question is what moves between tiers and who decides.
Not because it is faster. Because it removed a team, a skill set and a second network from the build sheet.