God why am I defending Amazon?
For reference transit cost has dropped over the years and is now about at $0.000247/GB or free if there is a peering agreement with the network the data is sent to.
This isn’t just AWS BTW; all tier 1 cloud providers recoup their costs this way.
And kindly refrain from accusing others of Stockholm Syndrome here. It’s incredibly rude.
Finding any excuse to justify a 100x profit margin is a form of Stockholm syndrome.
We are comparing to T1 ISPs because they also provide huge amounts of bandwidth reliably, and they do it 3 orders of magnitude cheaper. Therefore it cannot possibly be the case that providing the bandwidth is simply that expensive.
Many of us are aware that ISPs charge less for bandwidth. But they are not offering the same thing as a tier 1 cloud provider.
You may find these interesting:
https://aws.amazon.com/video/watch/c37546e1558/
https://cloud.google.com/blog/products/networking/speed-scal...
Again, these egress fees are not paying for egress alone. They are defraying the cost of operating a massive, complicated, low-latency, reliable cloud network infrastructure, much of which is free of charge when used internally.
Some, like you, are expecting that cloud providers operate on a “cost plus” model where what you pay is strictly based on what marginal costs for the same thing. But it’s not that simple in reality. Charges levied for one thing can be used to pay for another.
If it's surplus it's not costs. Why did you claim that the high price is to cover the costs if 99.9% of it isn't covering the costs?
If I may take the opposite view: companies should be grateful to get any profit from me at all.
I edited my comment above to explain in more detail what these charges are covering (and it’s not egress bandwidth alone). This has been a fast moving conversation and I’ve tried to add more context, unfortunately after the fact. My bad for posting too quickly.
They could charge 10x-ish as much as T1 ISPs and cover all those costs and still have it be almost all profit.
Yes, you're legally correct, but the point is AWS are clearly overcharging for egress.
As you say:
> Egress pricing pays for the AWS network infrastructure that is incredibly reliable (hardware failures happen regularly yet almost no one notices) and allows it to operate at scale sufficient to absorb even the largest DDoS attacks.
Cloudflare is also pretty good at absorbing DDoS attacks, yet charge no egress costs.
Maybe the real question is, are the tier-1 providers colluding in overcharging for egress?
Egress costs are fairly clearly a vendor lock in strategy. If I want to transfer 1TB to run on cheaper compute at another provider, that costs me $90 (edit) on AWS as opposed to $0.24 at wholesale prices. It's unlikely to be cost effective so I choose AWS compute.
EDIT: thanks to someonebaggy for pointing out my decimal point error.
Ninety dollars to transfer a terabyte is clearly ridiculous.
Many people avoid T1 clouds because they're so expensive, but they seem to have found an extremely profitable market segmentation consisting of the remainder.
There's only a small little mention, if we are lucky, on the couple drives that have it (expensive enterprise flagships). It should be a regular sticking point, whether it's there or not. Without pressure it's not going to get regularly available, it feels like.
FDP is so simple. Declare a number for what pool of data you want to write into. Data of the same pool gets written to the same storage such that you can wipe it latter together. It has huge wins though against write amplification! Massive wins. For so close to free.
Some day I want to own a FDP drive. And then I can finally start using the tokio/io-uring support that I contributed! https://github.com/tokio-rs/io-uring/issues/380
The NVMe-KV is more radical. Still worth putting some pressure on, but your drive as KV, as object store, feels harder. Side note, really enjoyed this ceph nvme-kv offload post thing, my favorite tech write up in a while! https://ceph.io/en/news/blog/2026/for-whom-the-door-bell-tol...
FDP is a very minimal addition of control. But yes it is another box to ticket, is another place to upcharge. Yet still, we're only just seeing mainstream products emerge. Kioxia's CM10 for example. https://www.techpowerup.com/351218/kioxia-introduces-first-p...
I'm hoping that the need is great enough to break the industry control. I think for a while the market felt relatively well enough served such that it was unclear whether clear wins would actually result in customers. With AI need for speed, I can definitely imagine incredibly crazy CXL controllers or what not, that allow low latency acces to many many open channel flash systems. Wouldn't that be a thing.
Again though, FDP is such a ridiculously tiny add, and it helps SSDs so much for so many use cases. I really hope it becomes an expectation, not a feature, sooner rather than latter.
Edit: happy to see a new group has shown up asking for open channel flash, Open Flash Platform. https://openflashplatform.org/
But also while it should be faster than it is, how would an S3 designed around SSDs look differently to an API user? I would think the API is basically the same.
- Directory entries are no longer returned in sorted order in ListObjectsV2
- There's an AppendObject API
- There's a RenameObject API
I suspect it's more to do with the fact that with One Zone is a clean rewrite of large parts of the application stack that makes up S3.
S3 is made up of hundreds of microservices[0], there probably isn't anyone at Amazon that actually understands the whole system. Refactoring it to support these features probably requires coordination between a lot of different teams. They might have petabytes of metadata, making a change to how metadata is persisted probably requires a massive risky data migration.
[0]: "All in, S3 today is composed of hundreds of microservices" - https://www.allthingsdistributed.com/2023/07/building-and-op...
The standards say SFTP. Most everyone ignores that part and have been using S3 buckets by bilateral consensus instead for years.
It turned out that all of our peers supported blob storage better than SFTP, which has some “show stopper” problems like forced outages caused by mandatory host key rotations.
It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
S3 is terrible deal on any front. it's just easy
For one, this has no redundancy and doesn't factor the cost of the machine providing the storage. But also it's an upfront cost for storing 4TB. It would be implied you wouldn't literally store 4TB on the SSD because now you're out of capacity for more data. If you only have 2GB of data in a bucket then S3 is still magnitudes cheaper, faster, and more redundant than anything you could put together yourself because the cost is spread across all the users of the service.
Reality is there are so many times that S3 makes the most sense that it makes people short-circuit and always pick S3 despite the huge hidden IOPS and bandwidth costs everyone rightfully tries to point out.
Of course, that's list prices, you can get savings plans and RIs and discounts.
The cost of RAM per hour in AWS isn't cheap.
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
I'd bet the majority of projects, but perhaps not the majority of traffic, can run adequately on a single server and tolerate a few hour maintenance window one weekend every month. Don't waste money on overkill.
The biggest reason people pay for S3 is because of data transfer cost to other cloud services they are using.
Really? No, it's not a few percent.
You just plug it in and basically it runs. Thats pretty much it.
Having to explain to an HN audience how installing an SSD is trivial is weird.
Of course, if you had paid extra for the backups and redundancy, your data would survive.
So to address your original point, unless you are paying extra for redundancy, it is cheaper if you use a 4TB SSD at retail price as the original commenter discussed.
The parent commentator is under the illusion that AWS automatically means security and scale and reliability.
I was merely pointing out to him that Iranian attacks must serve as a wake up call for him.
Yes you can make the argument for that specific use case.
For everyday use case, using s3 as your primary storage is costly and is not at all ideal.
(Don't bother with Ceph Object Gateway unless you really really need it - since you're programming your own application, access the storage pool directly with librados)
If you do a clever bit of caching work in your app like we have with our apps Slyp and SlypBusiness, you could even make it insanely cheap.
We have a file manager that is ultimately responsible for images, files, videos and whatnot. It loads the images/files from the downloaded local cache when requested by a feature in the app.
A standard feature uploads and downloads using signed urls obtained from the backend service. The manager increments the file's version in the backend on each successful upload. For download, the manager compares the version against the locally cached version. If there is an update, the new file is downloaded in the background and overwrites the local cache and notifies all features using the file inside the app.
Cloudflare can shut you down and charge an arbitrarily high price if they don't like you and feel you're abusing their low prices.
Maybe maybe bringing their own IP would have solved the problem, but Cloudflare was obstructing everything and then did a cutoff without proper warning. It was really bad on Cloudflare's part.
They really hated that particular customer.
And you're going to have to pay $15 of storage costs for your terabyte.