Rendered at 18:29:13 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
bob1029 2 hours ago [-]
If you are using LLMs to interact with sites like GitLab and GitHub, and you have the option to use a GraphQL API, you should jump on it immediately.
GraphQL is absolutely terrible for human developers to interact with, but it's like Facebook could see into the future back in 2012. I cannot imagine a more perfect API surface for agents. With the REST API on GitHub, you can consume maybe 10 issue JSON blobs before your context window is blown out. With GraphQL constraining the results you can easily read hundreds in the same token budget.
Additionally, the # of requests your agents need to make can be reduced in many cases since GraphQL can join across types whereas REST APIs cannot. You essentially get savings in two dimensions here. Quota and raw token volume per logical response.
iamEAP 57 minutes ago [-]
I work closely with the team responsible for a large, self-hosted GitHub Enterprise instance. This is good advice for clients/consumers of GH data, but it can very easily lead to a lot of strain on the server-side. It’s not obvious what fields that you request are simple reads from tables or actually end up invoking git under the hood.
You could argue the rate limit guards should better reflect that, but that’s just not the reality of the system. Likely speaks to a lot of stability issues GitHub has been facing lately.
enormousness 50 minutes ago [-]
Isn't it kind of part and parcel of any GraphQL deployment to update it efficiently?
bob1029 34 minutes ago [-]
All of the queries my agents use select fields like issue title, body, labels, createdAt, updatedAt, etc. That's about it. I would hope that stuff is cached and efficient to read. I do not think GraphQL is a good way to interact with git. Running git on the CLI is the best way to interact with git.
dieselgate 35 minutes ago [-]
> It’s not obvious what fields that you request are simple reads from tables or actually end up invoking git under the hood.
I'm not familiar with graphql but what would make something "invoke git", is it a technical thing or hyperbole?
kccqzy 35 minutes ago [-]
That is true for almost any GraphQL backend not just GHE.
philipp-gayret 1 hours ago [-]
Although that may be true, the quota at least on GitHub depends on what you're quering. We worked directly with GitHub and a complex enough organisation hit with a GraphQL query can actually hit your hourly app limit before you even get a response.
As for GitLab, having hosted it for medium size organisations (~200 devs) and seeing how monorepo's work (they don't, we had GitLab's team show us that one page view made 50K db queries on our setup), please consult with your local admin team before firing GraphQL at it.
w29UiIm2Xz 22 minutes ago [-]
GraphQL is the AGENT=true of API responses.
aschobel 2 hours ago [-]
mine just is the gh cli. is the advantage of graphql that they can compose a query that would take multiple cli invocations?
mattkrick 1 hours ago [-]
CLI still has the possibility of being a little more token efficient, at worst it may use graphql behind the scenes.
For our company, we advertise the graphql schema to bots and they can one-shot whatever task they're trying to do. I've found that it's so good that I cancelled building an an MCP server and any skill. Just a well documented GQL schema. It's pretty remarkable.
agentdev001 23 minutes ago [-]
> "I've found that it's so good that I cancelled building an an MCP server and any skill. Just a well documented GQL schema. It's pretty remarkable."
This gave me a chuckle, there's some subtle irony here- especially if the documentation of the GQL schema was work that needed to be done!
enormousness 47 minutes ago [-]
I think you might also have a GQL schema that's particularly well suited to what those bots need. Either that or your backend schema is relatively simple and the GQL schema covers all its possible data compositions.
saasrivals 57 minutes ago [-]
[flagged]
sockbot 1 hours ago [-]
no the advantage of graphql is that the caller can limit the response to only the information that they need. in a REST API the caller has to filter out the extraneous information. using LLMs to do the filtering uses up context window
CharlieDigital 60 minutes ago [-]
This is not fully true.
Most agents will use curl | jq to slice what they need (assuming a known API)
agentdev001 21 minutes ago [-]
Yea- its just a matter of whether the work is delegated to the client or offered by the server. Making sure what is served is in a really quality schema is generally the most efficient path- ime
Providing kickbacks to the repos being scraped would be a good way to help fund open source projects and pay creators like streaming services do. Seems like they're headed in this direction - it would be a massive product differentiator over GH
aprentic 2 hours ago [-]
My first reaction was that I really like this idea.
If we had a system where people who access projects pay and popular FOSS developers get paid for it we'd have much better alignment.
My second thought was that bots would immediately try to circumvent such a plan. They'd probably spam Gitlab with fake repos to try to harvest those payouts.
cush 31 minutes ago [-]
> They'd probably spam Gitlab with fake repos to try to harvest those payouts.
Yeah I wonder if the math would shake out to make that make any sense. Each bot would require a paid subscription, so the only incentive for them to do this would be if there was some discoverability algorithm or SEO that that traffic helped push the content to real users
MeetingsBrowser 2 hours ago [-]
I like the idea, but that wasn’t my read.
The guidance given seems to hurt open source projects, not help.
> Make the project private if the traffic is not coming from the audience you built it for, which stops anonymous callers reaching it at all. Or upgrade to Premium or Ultimate for much higher limits.
latexr 2 hours ago [-]
Three minutes after kickbacks were announced there would be a flurry of new repos being created with bots repeatedly scraping them just to get those kickbacks.
anamexis 2 hours ago [-]
Presumably whoever is doing the scraping would need to pay, to get rate limits conducive to scraping.
cush 29 minutes ago [-]
Correct
sph 2 hours ago [-]
Also known as the Cobra effect
demibabs 2 hours ago [-]
Damn, we’re even having Claude write important press releases now
Jeremy1026 1 hours ago [-]
I put the text in to gptzero's AI checker and it came back with being Highly Confident it was 100% AI written. I don't think I've ever seen it that confident that the entire thing was AI before.
rcxdude 60 minutes ago [-]
Claude's style is obvious enough I don't think you really need the checker.
wiether 33 minutes ago [-]
My FF being in French, I was automatically served a French version and I just couldn't understand what I was reading. Each sentence unclear, no link between sentences...
svachalek 48 minutes ago [-]
Hmm I don't know if they've updated it, but I'm not smelling any of the Claudisms I know so well. No convoluted sentences, metaphors, therapy-speak, 'it's not' constructions, "load bearing" etc etc.
vips7L 2 hours ago [-]
sad days
zb3 1 hours ago [-]
"Write a post about decreasing limits so that this information is as obfuscated as possible."
rkagerer 1 hours ago [-]
60 requests per hour per IP if you haven't signed in... well that's unfortunately low.
ddtaylor 2 hours ago [-]
> A request that arrives with no credentials gets 60 requests per hour per IP address.
One request per minute.
mplanchard 2 hours ago [-]
Average, yes, but the way they phrase it, it could be a token bucket or similar, where you can do 60 quick requests and then be blocked for a bit while the bucket refills
Makes 12 graphql API calls and 2 /api/v4 calls... So you get like 4 pages per hour unauth?
sandeepkd 2 hours ago [-]
More like you have 60 requests, you can exhaust them in a single second or spread them differently as per your choice.
Jaxan 2 hours ago [-]
Doesn’t help with scrapers though. They use a unique IP for each query.
silverwind 1 hours ago [-]
Or a lot more with IPv6.
ddtaylor 1 hours ago [-]
Most connections or VPS have a shared 64bit prefix that acts similar to an IPv4 address in that it's easy to block.
Yes, you get 64 more bits to make whatever addresses you want, but the prefix is still your fingerprint.
MeetingsBrowser 2 hours ago [-]
Hopefully they bump this up.
Browsing open issues or reviewing a few PRs will easily use more than one request per minute.
The limits are based on the average user but I wonder if the most common interaction is to view a readme and bounce.
I don’t know that putting a paywall up to learn from or even consider contributing to public projects is a good thing.
tempest_ 3 hours ago [-]
I assume this is because of LLM scraping.
nijave 2 hours ago [-]
Presumably but I wish they'd also focus on optimizing code/making pages fully cacheable instead of rate-limits and blocking
swatcoder 2 hours ago [-]
Making requests is inherently cheaper than delivering responses, even with caches. Efficiency improvements can buy a little time on a given resources but won't solve the problem of bot saturation now that everybody can spawn a custom bot in about 12 seconds and is being encouraged to do so.
Rate limits, blocking, and pay-per-use are the only roads out and even those might not last as models get better at hacking and masquerading.
The internet we want to use LLM's with is simply not one that can support LLM's, and with LLM's not going anywhere, the whole experience of the internet is going to be forced into some radically less open and more expensive paradigm.
Policies like this just represent the beginning of the transition.
jayd16 2 hours ago [-]
The limits are for API requests, no? Or is this just an unrelated performance complaint?
nijave 13 minutes ago [-]
Last I checked, they just blanket added `cache-control: max-age=0` to everything.
We need a new non-commercial version of the internet. Free of bots, free of ads, ... you pay to access social media optimized to be interesting enough to be worth paying instead of addictive enough to keep you scrolling to show you more ads.
It may look impossible right now. But what is impossible for real is to continue as we are. The damage that internet does to society is increasing by the day while its value is reduced (economic value, social value).
godwinson__4-8 2 hours ago [-]
> We need a new non-commercial version of the internet.
> you pay to access social media optimized to be interesting...
So is it commercial or non commercial?
Nothing is stopping you from creating a social network that is pay gated. Go build it. If you can't get anyone to sign up perhaps you'll realize it's not so easy as scapegoating addictive social media.
There is a whole cottage industry of people who legitimately make their living criticizing Facebook. It's a consumer software product. Yet few of these people seem to have their conviction extend to building an alternative that ever catches an audience. Why is that? Because addiction? Any other excuses?
dwedge 57 minutes ago [-]
It only proves your point because it wasn't successful in the end, but I really enjoyed the experience on app.net (a paid version of Twitter) while it still existed. Personally I think they lost it where a lot of these alternative platforms lose it - they try to build an ecosystem that is always coming soon and ends up scattered instead of focusing on one product
nijave 12 minutes ago [-]
Yeah but you lose anonymity unless you can devise a scheme to identify advertisers and bots separately from and without identifying real humans.
phoe-krk 2 hours ago [-]
> We need a new non-commercial version of the internet.
You'd need a force strong enough to prevent it from falling prey to this tragedy of the commons, and that force would need to be stronger than the incentives to commercialize it. And that's where plenty of contemporary scraping-based salaries lay.
heavensteeth 2 hours ago [-]
That's this internet. Nothing precludes you from dropping big tech and solely engaging with the "indie web"[0][1] if you so choose. You're free to build your own website with its own RSS feed, join webrings[2] of like-minded people, and engage organically.
[0] One of many similar initiatives, I'm sure. Not an endorsement.
This is fine when you hide in the dark forest, small and unassuming, but the moment something discovers you, then you get ate by a predator.
Your human engagement will attract said predators because it's a unique information signal.
9dev 2 hours ago [-]
Careful what you wish for. The ticket to entry will (have to) be a proof of identity, the ultimate nail in the coffin of privacy and anonymity on the web.Outside of that, chaos.
pythonaut_16 57 minutes ago [-]
[dead]
immortalist 2 hours ago [-]
Impossible to create
armadyl 2 hours ago [-]
Also for better or for worse the ad-centric model to some extent has allowed more people to access information.
Gating everything behind paid (but with no ads) likely would hurt a significant amount of lower income users.
OtherShrezzing 2 hours ago [-]
I'm not so certain. I used to buy a broadsheet newspaper, which was full of ads. Now I pay a few hundred a year, and get the same newspaper online, with no ads at all.
So, the precursor to online media has already gone through this paradigm shift.
pixl97 2 hours ago [-]
I mean, what you're saying is "If I pay for a product ads go away"
Which is partially true, but it only shifts the distribution of the problem. Once your service gains enough popularity network effects cause it to gain value. You have to worry about high priced buyouts of the entire service (great for the site owner, terrible for the users).
Zambyte 2 hours ago [-]
It's not, it just requires creativity. One example I can think of is limiting connectivity by distance between nodes. Something like Meshtastic seems unlikely to ever be commercialized in the same way that the Internet has been. Sure, you would lose some useful applications of the Internet, but you would gain other things.
toomuchtodo 2 hours ago [-]
AT Protocol and PDS’. BitTorrent to share and distribute bundles of data.
lateatdesk 1 hours ago [-]
60 requests an hour per IP seems low for a school or office network. A few people browsing issues and source files could use that up quite fast.
serhack_ 2 hours ago [-]
I would spend thousands of dollars for gitlab in terms of:
1) better UX for admin panel, I'm not sure what I've enabled and what not. Several buttons do not disable the rest of the settings, leaving me with some doubts (e.g. if I disabled grafana, why is there a setting that talks about where/how I store?)
2) a minimal version of gitlab without all the AI
jtwaleson 2 hours ago [-]
I think it's because people are building agentic flows, reducing the amount of developer seats needed. It's the first step towards usage based pricing.
theokrueger 2 hours ago [-]
Github would hit four nines if they followed suit. no clue why the dont try
pkaye 2 hours ago [-]
Doesn't GitHub already have rate limiting especially if you are not logged in?
theokrueger 39 minutes ago [-]
it's too generous and contributes to their reliability woes
Jaxan 2 hours ago [-]
Yes. A lot is not accessible if not logged in.
shimman 1 hours ago [-]
Because that would go against MSFT's wishes on pushing LLM driven development if they started acknowledging that these tools are more damaging than helpful.
296012 2 hours ago [-]
Congrats on making the world worse with AI. All this performative data scraping and uploading and no progress at all.
ephemerally16c4 2 hours ago [-]
Making the privileged money and giving them the power to manipulate the mass is progress to some.
micromacrofoot 2 hours ago [-]
what are you talking about? it's progressing a lot of money into specific people's pockets
speedgoose 2 hours ago [-]
No progress?!
IhateAI_3 2 hours ago [-]
[dead]
2 hours ago [-]
bearjaws 2 hours ago [-]
I am honestly surprised they aren't going lower at this point.
Gitlab must pay a fortune to bot traffic, most of which is malicious or garbage at best.
throwitaway222 1 hours ago [-]
Walk back in 10...9...
mschuster91 46 minutes ago [-]
> You get 429 Too Many Requests with RateLimit-* headers and a Retry-After. Wait the interval it gives you, then retry.
Is there a test endpoint where one can validate the behavior of their ratelimit detection? Basically I do not want to cause excessive load on your servers just to test my implementation.
Retr0id 2 hours ago [-]
I understand why they're doing this, but the anticausative title kinda rubs me the wrong way.
MeetingsBrowser 2 hours ago [-]
Worst case, this could be the start of a paywall to learn from, contribute to, or host open source projects.
Hopefully they find some kind of carve out for OSS projects while still blocking the egregious offenders.
jmclnx 2 hours ago [-]
> The requested URL was not found on this server.
Getting that so I do not know exactly what they are doing. From the title I am guessing they are restricting or throttling if downloads exceeds some value.
sparkling 2 hours ago [-]
I noticed that recently Github.com has some kind of weird bot detection on public repos. I have a browser extension for switching User Agents for a specific legacy site, sometimes i forget to turn it off and Github will require me to login to view public repos.
All of this is most likely due to mass scraping by LLMs. Welcome to the total shitification of the web.
Macha 1 hours ago [-]
The post is about gitlab but on the subject of GitHub, the rate limit I seem to have for viewing commit history seems to be 0 for logged out users and viewing source files seems to be about 5/hour.
GraphQL is absolutely terrible for human developers to interact with, but it's like Facebook could see into the future back in 2012. I cannot imagine a more perfect API surface for agents. With the REST API on GitHub, you can consume maybe 10 issue JSON blobs before your context window is blown out. With GraphQL constraining the results you can easily read hundreds in the same token budget.
Additionally, the # of requests your agents need to make can be reduced in many cases since GraphQL can join across types whereas REST APIs cannot. You essentially get savings in two dimensions here. Quota and raw token volume per logical response.
You could argue the rate limit guards should better reflect that, but that’s just not the reality of the system. Likely speaks to a lot of stability issues GitHub has been facing lately.
I'm not familiar with graphql but what would make something "invoke git", is it a technical thing or hyperbole?
As for GitLab, having hosted it for medium size organisations (~200 devs) and seeing how monorepo's work (they don't, we had GitLab's team show us that one page view made 50K db queries on our setup), please consult with your local admin team before firing GraphQL at it.
For our company, we advertise the graphql schema to bots and they can one-shot whatever task they're trying to do. I've found that it's so good that I cancelled building an an MCP server and any skill. Just a well documented GQL schema. It's pretty remarkable.
This gave me a chuckle, there's some subtle irony here- especially if the documentation of the GQL schema was work that needed to be done!
Most agents will use curl | jq to slice what they need (assuming a known API)
If we had a system where people who access projects pay and popular FOSS developers get paid for it we'd have much better alignment.
My second thought was that bots would immediately try to circumvent such a plan. They'd probably spam Gitlab with fake repos to try to harvest those payouts.
Yeah I wonder if the math would shake out to make that make any sense. Each bot would require a paid subscription, so the only incentive for them to do this would be if there was some discoverability algorithm or SEO that that traffic helped push the content to real users
The guidance given seems to hurt open source projects, not help.
> Make the project private if the traffic is not coming from the audience you built it for, which stops anonymous callers reaching it at all. Or upgrade to Premium or Ultimate for much higher limits.
One request per minute.
https://gitlab.com/gitlab-org/gitlab
Makes 12 graphql API calls and 2 /api/v4 calls... So you get like 4 pages per hour unauth?
Yes, you get 64 more bits to make whatever addresses you want, but the prefix is still your fingerprint.
Browsing open issues or reviewing a few PRs will easily use more than one request per minute.
The limits are based on the average user but I wonder if the most common interaction is to view a readme and bounce.
I don’t know that putting a paywall up to learn from or even consider contributing to public projects is a good thing.
Rate limits, blocking, and pay-per-use are the only roads out and even those might not last as models get better at hacking and masquerading.
The internet we want to use LLM's with is simply not one that can support LLM's, and with LLM's not going anywhere, the whole experience of the internet is going to be forced into some radically less open and more expensive paradigm.
Policies like this just represent the beginning of the transition.
Loading https://gitlab.com/gitlab-org/gitlab is showing 14 API calls so presumably significantly more DB queries for a public page
It may look impossible right now. But what is impossible for real is to continue as we are. The damage that internet does to society is increasing by the day while its value is reduced (economic value, social value).
> you pay to access social media optimized to be interesting...
So is it commercial or non commercial?
Nothing is stopping you from creating a social network that is pay gated. Go build it. If you can't get anyone to sign up perhaps you'll realize it's not so easy as scapegoating addictive social media.
There is a whole cottage industry of people who legitimately make their living criticizing Facebook. It's a consumer software product. Yet few of these people seem to have their conviction extend to building an alternative that ever catches an audience. Why is that? Because addiction? Any other excuses?
Except it's non-commercial, therefore valuable, therefore commercialized, therefore commercial.
You'd need a force strong enough to prevent it from falling prey to this tragedy of the commons, and that force would need to be stronger than the incentives to commercialize it. And that's where plenty of contemporary scraping-based salaries lay.
[0] One of many similar initiatives, I'm sure. Not an endorsement.
[1] https://indieweb.org/
[2] https://en.wikipedia.org/wiki/Webring
Your human engagement will attract said predators because it's a unique information signal.
Gating everything behind paid (but with no ads) likely would hurt a significant amount of lower income users.
So, the precursor to online media has already gone through this paradigm shift.
Which is partially true, but it only shifts the distribution of the problem. Once your service gains enough popularity network effects cause it to gain value. You have to worry about high priced buyouts of the entire service (great for the site owner, terrible for the users).
Gitlab must pay a fortune to bot traffic, most of which is malicious or garbage at best.
Is there a test endpoint where one can validate the behavior of their ratelimit detection? Basically I do not want to cause excessive load on your servers just to test my implementation.
Hopefully they find some kind of carve out for OSS projects while still blocking the egregious offenders.
Getting that so I do not know exactly what they are doing. From the title I am guessing they are restricting or throttling if downloads exceeds some value.
All of this is most likely due to mass scraping by LLMs. Welcome to the total shitification of the web.