zlacker

[return to "Had a call with Reddit to discuss pricing"]
1. 58x14+Je[view] [source] 2023-05-31 18:31:34
>>robbie+(OP)
I think it's very clear that the recent LLM boom is directly responsible for Twitter, Reddit, and others quickly moving to restricted APIs with exorbitant pricing structures. I don't think these orgs really care much about third-party clients other than a nuisance consuming some fraction of their userbase.

Enterprise deals between these user generated content platforms and LLM platforms may well involve many billions of API requests, and the pricing is likely an order of magnitude less expensive per call due to the volume. The result is a cost-per-call that is cost-prohibitive at smaller scales, and undoubtedly the UGC platform operators are aware that they're pricing out third-party applications like Apollo and Pushshift. These operators need high baseline pricing so they can discount in negotiation with LLM clients.

Or, perhaps, it's the opposite: for instance, Reddit could be developing its own first-party language model, and any other model with access to semi-realtime data is a potentially existential competitor. The best strategic route is to make it economically infeasible for some hypothetical competitor to arise, while still generating revenue from clients willing to pay these much higher rates.

Ultimately, this seems to be playing out as the endgame of the open internet v. corporate consolidation, and while it's unclear who's winning, I think it's pretty obvious that most of us are losing.

◧◩
2. eru+Zt1[view] [source] 2023-06-01 01:58:32
>>58x14+Je
If you want training data for an LLM and are actively talking to some data providers, you'd just ask for a dump, instead of making a billion small requests.

(You'd make the billion small requests, if you are doing this on the sly.)

◧◩◪
3. fennec+ZW2[view] [source] 2023-06-01 15:23:08
>>eru+Zt1
Eh can still just automate creating a bunch of accounts and do it manually. Use one of the many captcha completion services where you pay for people to complete captchas for you. ML models can already pretty much do them anyway.

Then rotate between accounts and put a random time between requests. Restrict certain accounts to browse within certain hours/timezones. Load pages as usual and just scrape data from the page rather than via api.

However, I believe in a company's right to charge whatever they want for their services. But I also believe in the right for people to choose not to use that service and for freer alternatives to spring up.

Just like Tumblr, Reddit seem intent on killing themselves, although these days I'm not so sure. When Elon took over Twitter everyone was saying that all the users would leave and it would die. This is not the case, human nature means that people seek familiarity and will cling on, hmm.

[go to top]