TLDR
Tarotdoxa is an AI tarot reading app for iPhone and Android. It’s its own company, separate from my agency, and I’ve been building it with AI since March 15. The first session had a deadline in it: get into the App Store by March 27. It’s September. The app has four separate codebases, 539 iOS unit tests, a reading voice built from a real reader’s teaching, two Apple rejections, a Google approval it’s still sitting on, and a build I pulled out of Apple’s review queue myself on September 22. This is the honest version of how building an app with AI goes, including the night my own app locked me out at a party.
Where It Started
March 15, around 9 PM. The first working session’s context line said: “App Store deadline March 27 for Ghost Conference.” The idea was to have the app live in time for a conference two weeks later.
That deadline missed by a little.
The first night went great, which is part of the trap. By the end of it, the reading backend was live and verified, a privacy policy was deployed, and everything was committed, with no corrections needed from me. When a first session goes that smoothly, you start doing math in your head about how fast the rest will go. Don’t do that math.
Day two set the most important rule of the whole project. The first version of the reader’s system prompt used a generic voice: a wise reader with “20 years of practice.” It read like every other AI tarot app. I flagged it as not enough, and it took six more prompt iterations to fix, because we were polishing the wrong thing.
The fix was to stop inventing a voice and extract a real one. We pulled the voice from a working tarot reader’s recorded teaching, 19 transcripts and 5 case-study essays: how she nicknames cards, how she teaches, how she frames a hard card as “homework from the universe.” The lesson went into the project file in capital letters, more or less: “Never start with a generic voice when a real person’s content exists. Extract from their actual transcripts/essays FIRST, then build the prompt.”
If your AI product has a personality, somebody real should be underneath it.
Four Apps Wearing One Name
From the outside, Tarotdoxa is one app. Under the hood it’s four products that have to agree with each other:
- The iPhone app, written in Swift and SwiftUI
- The Android app, written natively in Kotlin and Jetpack Compose. Not a port. Its own code, matched to the iPhone design
- The website and web reading room, built in Astro on Cloudflare
- The reading engine, a Cloudflare Worker that owns the prompts, the card library, and every AI call
The engine routes work to different models on purpose. Claude Opus 5 does the actual readings, after a bake-off. A different model writes the astrology copy. A third sits in the failover chain in case the main model is down, and a small cheap model writes the one-line push notification teaser.
Four surfaces is where the pain lives. On August 10, a fix to one card’s meaning shipped to part of the stack and not the rest, and nothing caught it. What I said that day is still true: “Anything that we are fixing like this will, more often than not, need to be fixed throughout every level.”
Now an automated alignment check compares all 78 cards across all four surfaces on every deploy and again every two hours. If the iPhone and the website disagree about what a card means, I know before a user does.
The Randomness Is the Product
Two early rulings shaped how this app gets tested, and both surprised the AI.
The first: no rigged readings, ever. The normal way to test software is to force an input and check the output. For a tarot app that means “force the Tower card and see what happens.” I said no: “never muddy things with deterministic readings where we know what the output is going to be. They should always truly be fully random.” The shuffle is the product. If the testing process can fake a draw, eventually a fake draw ships.
The second was less philosophical. The AI kept misspelling the app’s own name in its documents. My note back: “you spelled it the same way you would spell pterodactyl.” That one’s in the lessons file too.
The design follows the same idea. We call the visual system sgraffito, after the technique where you scratch through a dark top layer to reveal the color underneath. The brand file says it better than I can: “You don’t paint meaning onto a situation, you scrape away the obscuring layer to find what was already underneath. The brand mechanic IS the product.”
What It Costs to Read One Card
AI apps have a cost per use that normal apps don’t. Every reading is a real bill.
At the end of July I switched the reader to Claude Opus 5, and the first measurement said reading costs had doubled. That would have been a big problem. It was also wrong.
The first analysis split the data by date, before the switch and after. But other things changed in that window too. When we re-split it by which model each reading was sent to, the real increase was somewhere between 1.18x and 1.58x, depending on the reading. The rule we wrote down: “When splitting a cost series by ‘before and after a change’, split on the thing that changed, never on the date.”
The bigger lesson was where the money actually goes. Every reading sends a long, fixed instruction set, around 26,000 to 27,000 tokens, before the user’s question even starts. When that instruction set has to be sent cold, it’s about seven-eighths of the cost of the whole reading. When it’s cached, it’s a fraction of that. A cold reading can cost roughly ten times a warm one.
So one of the most valuable pieces of code in the whole app does nothing a user would ever see. It keeps the cache warm.
Launch Night
August 31. Everything had to go in the same night.
The team’s own note about that night is the best ship-night advice I’ve got: “On a ship night, find the step with a vendor-controlled wait and start it first. Every other task is then free.” App review is the slow step you don’t control, so everything else gets arranged around starting it.
Three things stood out.
The checker blocked me three times, and it was right. My AI code review gate blocked the release three rounds in a row, at 96%, 96%, and 98% confidence, over the audio narration feature. It had a real race condition where a refund and a re-grant could loop, and one bug fixed in only one of the two places it lived. Instead of patching it a fourth time at midnight, I turned the feature off and shipped without it. It came back the day after submission, once the fixes held up under a full test pass.
The screenshots failed with no explanation. Apple’s upload system rejected the screenshots five times with a generic server error. The real cause, found by reading the image files’ headers directly, was two bad files: one saved at 16-bit color depth and one with a transparency channel. Neither is allowed. The error message said neither.
At 10:16 PM, iOS 1.0 went into review with 18 items attached: the app version, the subscriptions, and the in-app purchases. The Android version went to Google the same night. My message when it went through: “wooohoooo! onto google play.”
Two Rejections, and What the Second One Taught Me
September 1: rejected under Guideline 3.1.2. The app description mentioned a free trial but didn’t state its length, the price after it, or link the terms and privacy policy right there. Fair. Fixed and resubmitted the same day, no new build needed.
September 4: rejected under Guideline 4.3(b), “spam.” Apple said the app duplicates the content and functionality of similar fortune-telling and astrology apps. In other words: there are already enough of these.
Here’s the part that stung. When we pulled the review notes Apple had actually read, straight from their system instead of from memory, we found our argument for why Tarotdoxa is different had never made it into the submission. The notes only answered the first rejection. We’d been so focused on the last problem that we never made the case for the new one.
The lesson, word for word from the project file: “a rejection response that only answers the LAST rejection leaves the new one unanswered. Check what the submission actually says before assuming the argument was made.”
We checked the landscape too: 21 tarot apps had been first released on the US App Store in 2026 alone. The rule isn’t “no tarot apps.” It’s “show me why yours isn’t another one.”
The fix that morning was all presentation, zero code: rewritten review notes, keywords, subtitle, and description. Resubmitted by 8:52 AM. That afternoon we set up a reviewer test account with real history in it (14 readings, 3 chat threads, 5 notes, and 2 audio readings) so the reviewer would see what the app actually does instead of an empty new account, and built a 29-second preview video straight from the simulator.
The Night My Own App Locked Me Out
September 12, at a party. I pulled out my phone to show someone the app, and hit my own paywall. So did one of our founding testers standing next to me.
Not a bug, exactly. Two weeks earlier we’d deliberately turned off the unlimited-access flag on our own accounts, so our trials ran out like any real user’s. Which is exactly what a real user would hit.
Then it got worse on the other phone. The paywall’s only exit was a “Sign out” button, and it signed out completely, with no visible way back in, because that screen reused the trial onboarding screen, which has no sign-in option.
No automated test would ever have found that. It took two people using their own app on a Saturday night. We fixed the dead end the next day, and that turned into one of the busiest days of the project: three new iPhone builds in one day. One fixed the lockout. One changed how the trial works. One added a backup way to pay on the web if an Apple purchase fails, which is now allowed for US users.
The trial change came straight from me: “the way I want the 7-day trial to work is more like a 7-day money-back guarantee. If they sign up and they’ve agreed, once the 7th day hits, that’s when the charge goes through.” That afternoon the Android version with credit packs got approved by Google too.
Why I Pulled It Myself
The first time I pulled the app from review, the build waiting in Apple’s queue still described the old trial, which didn’t match the product anymore. I pulled it and resubmitted the one that matched.
Then on September 22, I pulled it again. Not because anything was wrong with it. Because I think it can be better, and I’d rather launch the better version than get approved on the good one. Here’s what I told my AI team: “I feel like the app could just be more vibrant, and we don’t have to go so hard on the chalkboard-scraping-away-what’s-underneath vibe.”
That’s a real design pass, not a tweak. It’s also the strongest possible answer to “there are already enough of these apps.”
Google approved the Android version twice. It’s sitting there, approved and unpublished, because of a rule I set on September 2: both stores launch together. I’m not going to launch on one platform and make everyone on the other one wait.
So as of today: two Apple rejections, two withdrawals by me, zero public launches. I know how that sounds. I also know the version that launches will be the one I actually want people to open.
What AI Was Great At, and What It Wasn’t
Great at: speed across whole arcs. On day two, the part of the app that connects cards to each other in a spread, elemental pairings, number echoes, what a reversal is asking of you, went from research to deployed, tested code, and was shared with another project, in one sitting. That’s the kind of thing that used to be a week.
Great at: holding four codebases in its head. Nobody on a team of one keeps an iPhone app, a native Android app, a website, and an engine in sync alone. With the right checks, the AI can.
Not great at: fixing the class, not just the symptom. The pattern that shows up again and again in my notes: a fix addresses the reported problem, a second check finds the same kind of problem somewhere else, and only the third or fourth pass closes it. The worst one was a purchase check on Android that got “hardened” three times in a row, and each time the hardening introduced a new way to throw away a purchase the customer had already paid for. The fix that finally held asked Google directly what the purchase was, instead of trusting the app’s description of it.
Not great at: noticing when a number is lying. For months, the logs recorded which model answered each reading. It turned out the label was hardcoded, so it said the same model name no matter which model actually answered. A comment in the code even said the backup model “has never fired,” which was reading a constant, not a measurement. Fixed September 5. Anything concluded from the old labels had to be treated as unproven.
And my favorite: improving a fallback deleted a bug’s warning sign. Cards without custom content used to fall back to old public-domain text so generic that a missing card jumped off the page. When we upgraded the fallback to sound like our voice, the gap became invisible. The same change that made the app better made a missing card impossible to spot by reading. That’s why the alignment check exists.
The Numbers
- 169 days from the first commit to the first App Store submission
- 4 codebases: iPhone, Android, web, and the reading engine
- 539 iOS unit tests on the September builds
- 78 cards checked across all 4 surfaces on every deploy
- About 29 TestFlight builds between August 21 and September 13
- 2 Apple rejections, 2 Google approvals, 2 withdrawals by me
- 0 public launches, for now
What I’d Tell You Before You Build an App With AI
- Build on a real person’s knowledge. A generic AI voice is a commodity. Somebody’s actual expertise is a moat.
- Count your surfaces on day one. Every platform you add multiplies every fix. Build the check that keeps them in sync before you need it.
- Measure cost per use from the first reading. And when a number jumps, split the data on the thing that changed.
- Assume the review notes are what gets read, not what’s in your head. Read back exactly what you submitted.
- Use your own app like a stranger. Turn off your own special access. Your worst bugs are in the flows only a real user walks through.
- Let the checker block you. Shipping without a feature beats shipping a known bug at midnight.
- It’s okay to launch later. Nobody remembers when you launched. They remember what they opened.
I appreciate you reading along on this one. It’s been the most humbling build of the year, and I’m not done. If you’re thinking about building an app, or you want AI working inside the one you already have, book a call from the home page. And if you want to see how I keep a build this size from going sideways, the harness post is the other half of this story.
Build first, learn fast.
Keep reading
- The Ghost Conference Website That Became an Operating System (Haunted Nav Bar Included)The old site timed out. Six months and 1,421 commits later, the Oregon Ghost Conference runs on two websites: a haunted public site with its own checkout, and Mission Control, a private back office that approves applications, builds every profile page, and publishes the public site with one click, with the inbox and the door register ready for conference weekend. Here's every room and every pipe.
- How I Built a Full Website in 30 Hours (After Failing Spectacularly the First Time)I spent two weeks planning a website with Claude before writing a single line of code. Then I built it in 30 hours across one weekend. But the real story isn't the build weekend - it's the complete failure that forced me to rethink how I work with AI from the ground up.
- I Let an AI Drive Pro Tools for an Entire Record (The Best Thing It Did Was Put Things Back)Eleven mixes, two days, one AI model driving real Pro Tools with a mouse and a scripting bridge. It fixed a mix that was clipping at +9 dBTP, trimmed every song to a shared headroom target, tested at least fourteen tonal ideas, and put every one of them back. Here's exactly how it worked, and what it taught me about my own mixing.