For caching, only if you don't specify your preferred providers and let OpenRouter route each request itself. I have stuff like this in my OpenCode config for each model I use and I regularly get ~90-95% cache hit rates.
It still won't be quite as high as you'd get by just using DeepSeek because occasionally a request will fail and you'll get routed to a backup provider with nothing cached, but it's close enough not to matter in most instances.
But I can't argue with the lower off-peak pricing when using DeepSeek directly. The downside is they train their models on your input, which might be a deal-breaker for many users (as it is for me).
"Privacy#
All these models are hosted in the US. Providers follow a zero-retention policy and do not use your data for model training, with the following exceptions:
Big Pickle: During its free period, collected data may be used to improve the model.
DeepSeek V4 Flash Free: During its free period, collected data may be used to improve the model.
MiMo-V2.5 Free: During its free period, collected data may be used to improve the model.
Laguna S 2.1 Free: During its free period, collected data may be used to improve the model.
Ling-3.0-tiny Free: During its free period, collected data may be used to improve the model.
LongCat-2.0 Free: During its free period, collected data may be used to improve the model.
North Mini Code Free: During its free period, collected data may be retained and used to improve the model. Do not submit personal or confidential data. See the provider’s Terms of Use and Privacy Policy.
Nemotron 3 Ultra Free (NVIDIA free endpoints): Trial use only — do not submit personal or confidential data. Your use is logged for security purposes and to improve NVIDIA products and services. The logged session data for improvement purposes is not linked to your identity or any persistent identifier. For more information about data processing practices, see the Privacy Policy. By interacting with this endpoint, you consent to the collection, recording, and use of such information and the NVIDIA API Trial Terms of Service."
just checked, yes they don't seem to provide that information, most other providers are advertising fp8 or fp4 which is okay, but "together" doesn't, so they are likely using fp4
update: coreweave/fp8 is at 191 tps, launched this morning, but really bad cache hit rate (~60%), coreweave is good for privacy but let's hope they improve cache
I use it on fireworks which is US/ZDR and pretty reliable. We run a few hundred million tokens/day through it for dollars. Many are cached, which is super duper cheap.
Google did make location history private in the maps app(delete on server side) at some point, if I recall correctly everyone can at the least opt in(i may be wrong and it may even be opt out!). I believe out of the exact concerns of warrantless sweep of all location data of all users around a lat long for server side storage by law enforcement.
It went under the radar but is a good example of _some_ incremental positive improvement over centralized storage.
Per friends who have done this, you apparently need to explain why and have reasons™. It’s not just a quick “hi forget me” request. All of which I find a bit absurd.
I should have been more clear. You have to give reasons why, as in - if your identity is at risk, legal issues, company concerns..etc. This was the push back my friends and I have joked about (only two so far who have made this request).
The request is not taken at face value, you have to have it approved.
Welcome to the thunder dome. YC is not our friends, they just host what (was) a niche forum where early tech peeps could have mostly civil conversations. It’s become increasingly hostile in moderation and sharp political opinions, that a lot of the joy here has fumbled out. (sorry mod team..I know you deal with a lot and traffic/spam keeps growing).
It’s why the US has such a problem with gun control and laws.
I could argue everyone should have guns, and it’s just bad people, right?
reply