There Is No Formula for Ranking on ChatGPT
Every week someone sells a course, a plugin, or a consulting retainer promising to get you cited by ChatGPT, ranked in Gemini, or featured in Claude's answers. The pitch borrows SEO's old confidence: do these seven things, get this checklist, watch your brand climb the AI results. It is the same certainty that made SEO a $80 billion industry, applied to a system nobody selling the certainty actually understands.
I don't sell that certainty. I run byldr.dev as a live experiment in AI visibility, and the first thing worth saying is the uncomfortable one: nobody — not me, not OpenAI's own engineers half the time — can tell you with confidence why a large language model cites one business and not another for a given prompt. These are opaque, genuinely intelligent black boxes. Anyone promising a formula for cracking them open is either lying or hasn't tried hard enough to notice they're lying.
Why the black box is the whole point#
Traditional SEO had a legible target. Google published guidelines, the algorithm was reverse-engineerable through enough trial and error, and ranking factors could be tested in isolation because the system was, underneath everything, a deterministic ranking of indexed pages. You could build a formula because there was a formula to find.
LLMs don't work like that. The same prompt run twice can surface different sources. The model's citation behavior depends on training data cutoffs, retrieval layers bolted on after the fact, prompt phrasing, and weighting decisions that are proprietary and, in a lot of cases, not even fully understood by the teams that shipped them. GPT, Claude, and Gemini are three different black boxes trained on different data with different retrieval logic. A tactic that gets you cited in one may do nothing in another, or actively hurt you in a third. Treating this as a solvable formula is a category error. Treating it as an experiment is the only honest move left.
The experiment: picking a niche like it's a city#
The niche I picked is AI-driven email and order automation for businesses drowning in manual requests — the kind of company fielding hundreds of inbound emails a day that could be triaged, answered, or routed by AI instead of a person copy-pasting replies. It's specific enough that the buyer's questions are predictable, and narrow enough that I can actually saturate it instead of chasing a term as broad as 'AI automation,' which every vendor on earth is already fighting over.
Local SEO domination worked because you controlled every signal pointing at a defined territory. I'm doing the same thing here, except the territory is a subject area and the signals are citations, backlinks, and structured content that models can retrieve and quote back to a buyer asking a specific question. The geography changed. The domination logic didn't.
Reasoning backward from the buyer, not the model#
The mistake most GEO advice makes is starting with the model — what does GPT want, what does Claude reward, what schema markup does Gemini prefer. Start there and you're optimizing for a system you can't see inside. Start with the buyer instead. Someone drowning in manual email triage doesn't type 'best AI email automation platform.' They type something closer to what they'd say out loud to a colleague: 'how do I stop answering the same customer emails all day,' or 'AI tool to sort support tickets by urgency,' or 'automate order confirmation emails without hiring someone.'
Those real, messy, first-person questions are the actual surface area. If your content, your case studies, your product pages, and your third-party mentions aren't built around the language a real buyer uses under real pressure, it doesn't matter how well you've 'optimized for AI' — you're answering a question nobody's asking.
What tracking actually looks like#
Here's the part that replaces the fake formula with real discipline. Build a bank of roughly a thousand prompts — every plausible variation of the question a buyer in your niche might type into GPT, Claude, or Gemini when they're looking for something like what you sell. Not keywords. Full questions, in the buyer's voice, covering every angle: symptom-based, solution-based, comparison-based, budget-based.
Then run that bank across all three models and log whether your name, your product, or your business shows up in the answer — and how prominently. That's your baseline. It's not going to be flattering the first time. That's fine. The number itself matters less than what you do with it.
- 01Build the prompt bank
About a thousand real buyer questions in your topic niche, written the way a stressed-out human would actually type them.
- 02Run it across the models
Same prompt set through GPT, Claude, and Gemini. Log every citation, mention, and omission.
- 03Act on the gaps
Where you're absent, figure out what content, structure, or third-party signal is missing — and go fill it.
- 04Re-run monthly
Same prompts, same cadence. The trend line is the only signal worth trusting, not any single run.
Re-run the identical prompt set every month. Not a new set — the same one, so the comparison means something. If your visibility is climbing across the three models, the work is working. If it's flat or falling, you learn that fast enough to change course before you've sunk a quarter into a tactic that never had a mechanism behind it in the first place.
GEO is real. It's genuinely doable. It is not magic, and it is not a coin flip either — it's an experiment with a measurable outcome, run the way any competent engineer runs an experiment: define the input space, measure the output, change one thing, measure again. That's what I'm doing on byldr.dev with a niche instead of a zip code, and it's the only version of this work I'd put my name on.