I often hear: "Ceph is modern, SAN is old." That's a false debate.
In production, you don't choose a slogan. You choose an acceptable failure mode and a complexity level that the team can absorb at 3 AM.
What performance tests don't tell you
Both approaches can produce good numbers in nominal conditions. What matters is behavior in real situations:
- maintenance under load,
- partial network incident,
- sudden latency spike,
- lack of investigation capacity during a critical window.
That's where the differences appear.
Ceph: real power, mandatory discipline
Ceph can be excellent, but not "free" operationally.
It requires:
- a team capable of diagnosing quickly,
- solid observability,
- up-to-date runbooks,
- a rigorous change management framework.
Without that, complexity debt accumulates silently.
SAN: useful simplicity, dependency to accept
SAN can remain a very good choice when the priority is operational predictability.
The trade-off is known:
- less intrinsic flexibility,
- more vendor dependency,
- but often better operational resilience with a smaller team.
That trade-off can be perfectly rational.
What should guide the decision
- actual workload profile (not just averages),
- continuity requirements per service,
- current team capacity, not target capacity,
- governance and audit obligations.
When these elements are clear, the technical choice becomes almost self-evident.
Position
The best storage doesn't exist. The best storage for your context, however, does.
If the goal is stable production, prioritize the architecture the team can operate cleanly — not the one that impresses in a presentation.