M
Magdalena Śleboda
Head of Operations | Scaling Tech Organizations with AI, Automation & Data-Driven Execution | Global Ops & Transformation Leader
just now in The description of my XDS panel hosted by Summer O'Brien says policies, cost models, and risk frameworks are still catching up with the tools our teams already open every day. Accurate.
From the service provider perspective it’s no longer a question of if, but of how. Which model, which skills, how to optimize the toolkit and tokens you burn to get to the results you wanted. And who reviewed it before it went to the client’s codebase.
The model is the easy part. The harness around it is where the advantage sits. Craft problem, not a licensing one.
Max Wojczuk and Michal Niec have been publishing for months how Appliscale engineers experiment with it in internal and external projects - token usage, toolkit and agentic setup, what we tried and dropped. They go deeper than I’ll manage in a panel slot.
AI in External Development: Policy, Risk, and Reality. Vancouver, Sep 9, 11 am. I’m booking meetings around it if you want to chat more.
View on LinkedIn
M
Michał Nieć
CEO of Appliscale | AI in AdTech Expert | LP & Angel Investor in AdTech/GameTech
5 hours ago in tl;dr we're looking for Senior Data Engineer who can handle petabytes of data.
I write a lot about Ad Tech on LinkedIn, but that is only half the story at Appliscale .
The core of our DNA isn't advertising. It is scale. We build infrastructure for companies that generate data so fast it breaks normal cloud architectures.
Right now, we are hiring a Senior Data Engineer for one of our longest-standing partners: one of the largest game studios on earth, responsible for massive global MOBA and FPS franchises.
When millions of players log in concurrently to play an FPS, the data exhaust is staggering. You aren't just managing gigabytes of data; you are playing a vital role in building solutions capable of processing petabytes of information.
This role sits inside the Data Experiences and Automation team. The mission is to harness that massive data stream to power player-centric decisions, Machine Learning, and GenAI pipelines. You will be building the data foundation that insight analysts and ML engineers rely on to iterate fast.
We need a Senior Data Engineer who understands large-scale, globally distributed systems.
The details:
- B2B: 26,000 - 30,000 PLN net / mo.
- Remote / Krakow
- PST Overlap required (the core team is in LA and Seattle)
If you are tired of working on small datasets and want to manage petabyte-scale infrastructure for millions of players, you know the drill: https://www.appliscale.io/career/senior-data-engineer-pst-overlap-zbor
View on LinkedIn
Damian Naglak
Head of Engineering | Bedrock Platform | AdTech
1 day ago in You get the same ad ten times in one evening. That campaign almost certainly had a frequency cap set well below ten. A cap counts an identifier, and you keep arriving as a different identifier.
In the bidder a cap is a lookup. A request comes in, the bidder reads a counter for this person and decides whether to bid. For that number to mean anything, the request has to carry something that points at the same person as yesterday, and the impression from an hour ago has to be in the count already. Both of those fail often enough to explain the ten.
Start with the identifier. A bid request can carry a person in four ways, and any of them can be missing. The exchange's own ID for you (user.id), usually taken from its cookie, which it may rotate under its own privacy policy. Your own ID, mapped for you by the exchange (user.buyeruid), which exists only where a cookie sync ran first. The advertising ID from the phone or TV operating system (device.ifa), which anyone can reset from settings, and which arrives as a string of zeros on iOS when the user declines tracking. And IDs from third party graphs like RampID, ID5 or UID2 (user.eids), which arrive encrypted and encoded for one specific buyer, readable only with an agreement and the keys.
Those graph IDs carry a note saying how the match was made. Some come from an email login. Some come from a cookie sync. One of the values means the match was inferred from IP address and user agent, so the ID is a guess about who this is. A cap counts against a guessed ID the same way it counts against a real one.
When none of them arrive, what is left is the IP address and the user agent string (device.ip, device.ua). An IP is rarely one person. A home router puts the whole family behind one address. Mobile operators run carrier grade NAT, where thousands of subscribers share a single public address. The address does not hold still either: a home lease runs from a day to a week and can change when it renews or the router restarts. The user agent stopped separating people too, because every Chrome on Android now reports the same made up device.
Then there is the delay. The counter moves when the impression is booked as spend, which happens after the ad renders, and rendering can wait. A video prefetched during a game level plays when the level ends. An ad stitched into a podcast sits in the file until someone reaches that minute. On the web that gap is short. On formats that get cached or stitched into a stream it runs far longer. Everything you bid in that window is decided on the old count.
A frequency cap is worth no more than the identifier it hangs on. The harder it gets to recognise a person, the more of the same ad that person potentially sees.
Bedrock Platform
View on LinkedIn
Maksymilian Wojczuk
Technical Engineering Manager @Appliscale | Co-founder @DiPA
1 day ago in On August 17, Activision reported a connectivity incident affecting multiple Call of Duty titles.
The incident started at 6:43 PM UTC, with players reporting problems connecting to the game and logging in. Activision marked it as under investigation. I haven't found a public post-mortem explaining what caused it yet.
That leaves an interesting question for anyone working on game backends: where did it actually fail?
A player pressing “Play” can trigger a surprisingly long chain of backend services:
authentication → verification → entitlements → player profile → matchmaking → session allocation → game server
A failure in any of these can prevent the player from getting into a game, even when the game servers themselves are healthy.
This is where game backend engineering starts looking very familiar to anyone who has worked on distributed systems. You deal with partial failures, retries, timeouts, inconsistent state, traffic spikes, queues and regional failures. The player just experiences it as “the game doesn't work.”
This is also the area my team at Appliscale specializes in. We build backend services for authentication, verification, player preferences, queueing, and session management at large scale.
At scale, game technology is distributed systems engineering. Authentication, session management and player state are all part of the system the player depends on before the game even starts.
View on LinkedIn
M
Michał Nieć
CEO of Appliscale | AI in AdTech Expert | LP & Angel Investor in AdTech/GameTech
2 days ago in Really impressive AI parkour simulation by NVIDIA . The coolest part: all of that generated from just 30 seconds of reference material.
But looking under the hood, I think there's a fascinating tension robotics optimization vs. game realism I've been thinking about:
Robotics = optimization is king. We treat movement as a pure kinematic problem ("get from A->B via path C") to strictly minimize energy costs and wear-and-tear. The goal is always maximum efficiency.
However, to avoid the uncanny valley in characters, we actually need to inject somę “inefficiency”. If an AI/NPC moves with perfect thermodynamic optimization (like a robot), it feels… robotic. Humans don't move optimally - our actuators fatigue, joints stiffen, we run out of breath etc.
So to mirror human movement, we can't just optimize for the "best path" but rather “simulate the limits”: tired joints that fail under load, momentum that naturally drains speed, and recovery phases where efficiency drops.
For robots it’s ok to treat energy as infinite because they lack an explicit "battery" state or metabolic decay in their loss functions but that’s a different story for games.
(strongly recommed to follow Maksymilian Wojczuk from our Appliscale team who shares his deep dives into AI/gametech)
Paper: https://jiashunwang.github.io/HIL/static/mat/Hybrid_Imitation_Learning_TOG.pdf
Video: https://www.youtube.com/watch?v=8B05cy3UuSE
View on LinkedIn
Damian Naglak
Head of Engineering | Bedrock Platform | AdTech
3 days ago in A sportsbook asks for the people most likely to bet on football. The closest match anywhere in the data lives in a state where the company holds no licence. The model was right and the answer is worthless.
The last two posts were about finding the closest people among billions. Nobody wrote down who belongs in the audience and no score makes somebody a member, so everybody scores something against the description. Comparing against everybody is too slow, so each person is stored next to a short list of their nearest neighbours, and the search hops from person to person, always towards somebody closer. A few hundred stops instead of a billion comparisons. What comes back is a ranking, and it says nothing about who you may advertise to.
A profile often says where someone lives, so the vector half knows. Pooling flattens a whole browsing history into one vector, so two words about a city sit among hundreds of others, and one state reads much like the next. The vector leans one way, it never certifies. And the conditions that disqualify a person were never in their browsing history at all. Whether the sportsbook holds a licence in that state sits in the company's records, and changes when a legislature votes. Whether they agreed to be advertised to sits in a consent record. Whether you can reach them depends on an ID you may bid on. Anyone who already opened an account is on the advertiser's exclusion list. A better model ranks people better and learns none of this.
The obvious way to combine them is to take the closest hundred and drop the ones you cannot use, which in a narrow market leaves nobody. So the licence check runs during the hopping instead. Every time the search reaches somebody it asks two things: are they close, and are they allowed. Only the allowed ones go on the shortlist, and the rest still get hopped over, because their neighbour lists are the route to everybody behind them. The search stops when nobody left to visit is closer than the worst person on the shortlist, so when few qualify it fills slowly and the search keeps running. If one person in twenty is allowed, the search stops at twenty of them for every usable name it collects. Removing a few people costs almost nothing. Requiring something rare is what stretches the search.
The way out is to split people in advance by a separate index per state, so everybody inside it is eligible. Licences, caps, budgets and exclusions change by the second, so those get checked as the search runs. When very few qualify, the better move is to stop hopping and compare against the eligible ones directly, which is what engines do below a size threshold. And the top of the ranking wears away as a campaign runs: the best matches get served first, so they cap out first, and the ones who sign up move to the exclusion list.
Search over vectors is quick because it never looks at most of the people. Every rule about who you may advertise to is a reason to look anyway.
Bedrock Platform
View on LinkedIn
M
Michał Nieć
CEO of Appliscale | AI in AdTech Expert | LP & Angel Investor in AdTech/GameTech
3 days ago in A cheaper and more efficient that Seedance 2.5 flow to generate game videos.
Instead of prompting a video model directly and relying on multiple rerolls to get the right shot, GMI Cloud used DeepSeek V4 Flash to write a raw Three.js scene (exact camera movement, character motion, timing etc.).
Once the sequence was correct, they fed that raw, low-poly footage into MiniMax H3 to generate the final realistic render in a single take.
Cost: 48 mins, $1.97 total, one H3 generation.
View on LinkedIn
Damian Naglak
Head of Engineering | Bedrock Platform | AdTech
1 week ago in Embedding targeting is usually explained as comparing two vectors. The step before that, picking which vectors to compare against at all, is where most of the work sits.
Last post: you cannot put vectors in order, and you cannot rule out a group without looking inside it. So whatever makes this fast has to be built before any request arrives. There are thirty years of methods for doing that. Trees, which cut the space into smaller and smaller boxes, stopped working above about ten numbers per vector. Hashing, which sorts vectors into buckets using random cuts, needs too many buckets to get accurate. Grouping, which clusters everything and opens only the nearest few clusters, survives underneath the others. And one approach won outright, a graph of neighbours called HNSW, for hierarchical navigable small world. Of the ready-made databases sold for storing and searching vectors, Qdrant offers nothing else at all, and Weaviate makes it the default.
It works like walking downhill in fog. You cannot see the valley, but from where you stand you can feel which way the ground drops, so you step that way and check again. You stop when every direction goes up. The structure is exactly that. Every vector gets a list of about sixteen of its nearest neighbours, stored right beside it, and that list is all it can see from where it stands. A request drops in somewhere, checks those neighbours, moves to whichever one sits closer to what it is looking for, and does the same again from there, until none of them is closer than where it already is. Then it stops. A few hundred steps out of a million, and everything it never walked past is never seen.
There is one shortcut on top of that. The map comes in layers: a sparse top layer where the few points sit far apart and every hop covers a lot of ground, then denser layers below with shorter hops. A request crosses the rough map first to get into the right region, then drops down and refines. Flights, then trains, then walking.
What goes wrong comes out of the same picture. You can end up at the bottom of a small dip with the ground rising in every direction, while the real valley sits over the ridge you never crossed. The fix is to walk several routes at once instead of one, and how many is a configurable number. Walk more of them and you land in the real valley more often, and it takes longer. None of this is free. Those sixteen links per vector get stored alongside them, so the structure takes up more room than the data it was built from, and assembling it for a million takes time. You are buying query time with space, paid once in advance, and in bidding query time is the one thing you cannot get more of.
That trade is what makes matching by meaning possible inside an auction at all. A bid response has about a hundred milliseconds. Comparing one vector against every stored one does not fit inside that but walking to a few hundred of them does.
Bedrock Platform
View on LinkedIn
Damian Naglak
Head of Engineering | Bedrock Platform | AdTech
1 week ago in Segment targeting asks one question: is this user in list 12345. Embedding targeting asks which of my audiences this user is closest to. One of those is a lookup. The other is a search.
With segments somebody drew the line in advance, so a user is on the list or off it and the auction only has to check. With vectors there is no list and no line. Every user has some score against every audience, nothing is in or out, and the only way to decide is to rank everything and take the top few. So matching on meaning means finding the closest, and finding the closest is a different job from checking membership. Databases are fast on huge tables for two reasons. The rows are kept in order, so finding one is jumping straight to it. And they are grouped into chunks whose contents are known, so most chunks never get read at all: if a chunk holds only values between 10 and 20 and you asked for 50, it is thrown away untouched. A billion rows, and the database reads almost none of them.
Neither trick works on vectors. Keeping things in order needs something to order them by, and closest has nothing, because it depends on a question that arrives later and is different every time. Whatever order you pick, the next request wants another one. Skipping groups fails for its own reason: to throw a group away you have to be certain nothing inside it is closer than what you already found, and with hundreds of numbers per vector the distances all come out similar.
So you are left comparing against all of them. A million vectors at 3,072 numbers, four bytes each, is twelve gigabytes to read for one bid request. Nobody does that. Everyone gets around it the same way: group the vectors by closeness before any request arrives, then only look at the group the request lands in. That saves almost all the work, and it sometimes gives you the wrong answer. For bidding it lands somewhere specific. Today the hard part happened days earlier, when someone worked out who belongs in the audience and wrote it down, so the auction only reads the answer. Matching on meaning moves the finding into the auction itself, and what a bid costs stops being flat. Twice the audiences, twice the work, on every request.
Next: the three ways people organise vectors in advance, and what each one gets wrong.
Bedrock Platform
View on LinkedIn
M
Michał Nieć
CEO of Appliscale | AI in AdTech Expert | LP & Angel Investor in AdTech/GameTech
1 week ago in If you are not tired of AI content yet, congratulations, here is one more concept.
Agentic Software Factory. You stop writing code and start designing the factory that writes it. Nine stages, three kinds of actor. Agents implement, test, review and analyse incidents. Infrastructure sandboxes, ships and watches. Humans do four things: set intent, set priorities, own the architecture, and get paged when it breaks.
The pitch you will hear is that this removes the code review bottleneck. It does not. Ten agents opening forty pull requests an hour makes human review the bottleneck, harder than it has ever been.
What actually changes is where humans spend attention. Upstream, into the specification, because a vague spec used to cost you a day and now costs you forty confidently wrong diffs. And into the gates: tests, review agents, monitoring. Those are the only things that scale with the agents.
The factory does not remove the human. It removes the human from the middle and puts them at both ends.
So where is your bottleneck now: writing the code, or deciding what it should be?
View on LinkedIn
M
Michał Nieć
CEO of Appliscale | AI in AdTech Expert | LP & Angel Investor in AdTech/GameTech
1 week ago in I almost sent my team an update about a meeting that never happened.
I spent 30 minutes on that call with my friend Kamil talking about our content pipeline, in Polish, with Fireflies set to English by mistake. Out of half an hour it caught maybe ten sentences, mostly the English words that leak into any adtech conversation.
The transcript reads like this:
"Yes to die. Server CTV agency Programmatic in app advertising platform game advertising publisher data management. Varnish. What we shipped."
"No what good looks like. Just a call to action book the call. Solutions Boost TikTok Noitake cross platform cloud. No it is."
The summary built on top of it reads like this:
"AI Agent Development: Building scalable AI agents for programmatic in-app and streaming ads, integrating with publishers' data systems."
"Kamil highlighted the complexity of building AI agents in adtech, noting the challenge of aligning them with agency and publisher data management needs."
No he didn't, and every one of those bullets carries a timestamp pointing back at the noise it came from. With ten fragments and no way to say "I don't have enough here", the model simply wrote the meeting it assumed we'd had, and it happened to be a perfectly plausible one about the agentic advertising future.
I never opened that summary. Later the same day I asked Claude to pull the important updates from my client and partner meetings so I could keep the team in the loop, and the fake meeting made the list as one line about our PR process. I had my finger on send when the phrasing struck me as slightly off, and only then did I go digging.
That is what makes this worth writing about. The transcript is garbage that anyone would catch in a second, the summary is generic enough that you skim it and move on, and by the time it reaches a team update it is a single confident sentence nobody will ever question. Each layer launders the noise a little further from the evidence that would expose it.
View on LinkedIn
M
Michał Nieć
CEO of Appliscale | AI in AdTech Expert | LP & Angel Investor in AdTech/GameTech
2 weeks ago in The media narrative: "Gen Z is lazy and doesn't want to work"
The actual data: Gen Z is starting businesses at a rate that dwarfs every other generation alive.
I hire engineers every week and my experience with Gen Z is the exact opposite: they want to work hard, they just don't want to work on boring tasks that a machine should be doing and if a job is boring, they just start their own company.
We see this at Appliscale that the only way to attract and retain the best young talent right now is to offer them massive, complex challenges - based on my experience, more students crave for a job that require deep architectural thinking than a cushy repetitive job (anyway, the latter would soon be automated).
So if you offer them a boring job, they won't take it. They will just go build something better on their own which seems the right way.
PS. Yes, we are hiring still hiring! Check latest jobs: https://www.appliscale.io/career
View on LinkedIn
Damian Naglak
Head of Engineering | Bedrock Platform | AdTech
2 weeks ago in A page turns into a vector in 20 milliseconds. Eighteen times more text took 32 times longer, and asking the same card to read a 16,000 word history made it request 15 GiB in one block and give up.
Last week I timed one page and then priced it, both on page-length text. The long inputs are the interesting ones: a ninety day user history, a full CTV transcript. So this time I kept the model the same and made the text longer, on the same card. Long text costs more than its share. Take what 1,000 words cost: on the mid-sized model, 42ms inside a page, 40ms at 1,000 words, 44ms at 2,000, 54ms at 4,000, 76ms at 8,000. The 600M model went from 51ms per 1,000 words to 117ms at 32,000, where one text took 3.8 seconds. The same words cost nearly twice as much inside a long text as a short one.
A model does two jobs and their costs grow differently. Reading is the fair one: each word passes through the model's layers, so twice the words means twice the work. Comparing is the other: every word measures itself against every other word, which is how the model works out what each one means from those around it. Ten words means a hundred comparisons, a thousand words a million, so every time the text doubles, the comparing gets four times bigger. I opened a model up to check. Give it "the river bank was steep" and it builds a seven by seven grid of comparisons, every cell filled, twelve grids side by side at each of twelve layers. The comparisons run one way only: "steep" looking at "bank" scores 0.49, "bank" looking at "steep" scores 0.09.
On a page reading is the bigger cost, so the bill follows the length, as it did last week. Comparing overtakes it once the text passes about 768 words, this model's working width. The price per word climbed from 4,000 on. Memory suffers for the same reason: the whole grid of comparisons has to sit on the card at once. Sixteen thousand words squared, times sixteen grids, times four bytes a number, is 15.26 GiB in one lump on a card holding 22. Exactly what the crash asked for.
One way around it is to cut the text up and embed each piece on its own. The same 8,446 words cost 645ms in one go, 460ms as two halves, 374ms as four pieces and 336ms as eight. Half the price for the same words, and the result is a vector per session or scene instead of one for everything. It costs something: each piece is embedded blind to the others, so a reference in the last paragraph to something named in the first is lost.
Every model has a ceiling on how much text it will read. Getting near it costs more. The longer the text, the more every single word in it costs.
Bedrock Platform
View on LinkedIn
Maksymilian Wojczuk
Technical Engineering Manager @Appliscale | Co-founder @DiPA
2 weeks ago in "don't be a meat proxy" - 100% agree on this blog post from NIklas Gruhn
I can talk to Claude myself, don't need a human in the middle - copy-pasting a ticket to Claude, creating PR, pasting feedback back to Claude and repeat.
If you are just relaying prompt outputs, you are not the developer. The reviewer is - they are using the AI to build the system, using you as a proxy.
AI should amplify your intent, not remove it.
Link in comments.
View on LinkedIn
M
Michał Nieć
CEO of Appliscale | AI in AdTech Expert | LP & Angel Investor in AdTech/GameTech
2 weeks ago in We've been avid users of Airtable at Appliscale so the recent news about acquistion by Bending Spoons for $1.285 bln was quite a surprise.
Airtable raised a total of $1.35 bln in VC funding, hitting an $11.7 bln valuation at their peak in late 2021. Bending Spoons is acquiring them for $2.25 bln, but since Airtable is sitting on $965 mln in cash, the true enterprise value is just $1.285 bln. That’s a 90% drop.
Probably the most surprising is that this business is doing almost $500 mln ARR, and 20%+ growth which translates into 2.7x ARR. As an investor myself in couple of startups I understand that
a) VCs were likely pushing for liquidity and needed to cash out
b) The founders were probably massively diluted due to overfunding and simply wanted to take the exit and start fresh
Airtable used to be a huge unlock for us as an internal tool and a building block for our custom flows during no-code era, although with Claude/Codex its value proposition is definitely shifting. When AI can generate custom internal apps and scripts in seconds, paying a premium for a rigid no-code database probably makes less sense for some orgs.
So it makes you wonder: is this fire sale just a casualty of 2021 overvaluations, or another signal that AI is eating not only the no-code space but generally software?
View on LinkedIn
Damian Naglak
Head of Engineering | Bedrock Platform | AdTech
2 weeks ago in It's the middle of summer and IAB Tech Lab is shipping. AAMP 2.3 has no new deal types. It's plumbing: where the agents run, and who is allowed to claim they are an advertiser.
The agents now run where companies already run things, on Amazon Bedrock AgentCore and Databricks, with swappable storage so an agent is no longer stuck on one machine, and your own choice of model behind it. Google Ad Manager reporting and Meta buying got wired in.
Then a heading IAB called "trust and safety move from optional to enforceable," which covers two worries. The first is who you deal with: the buyer agent can now check a seller against the IAB Diligence Platform, a privacy questionnaire vendors fill in, before it buys. It's one yes or no, it comes off a list the buyer keeps itself, and it's off by default. The second is money. Earlier versions let the agent invent a CPM when it had no pricing data, and 2.3 removes that and records where every price came from.
On the seller side the matching question is who the buyer is. The seller agent has had buyer tiers since AAMP 2.0, and your tier sets the shape of the negotiation: a public buyer gets 3 rounds and at most 8 percent off the opening price, a seat 4 rounds and 12 percent, an agency 5 and 15, an advertiser 6 rounds and 20 percent. When the seller counters by splitting the difference, the public buyer gets 30 percent of that gap and the advertiser gets 65.
Until 16 July the tier came from the request body. Put an advertiser ID in and the pricing engine treated you as an advertiser, on quotes, counters, negotiation messages and template booking, with no ceiling applied anywhere. Twelve points of discount and three extra rounds, for filling in a field. Now the buyer passes the address of its agent card, the public profile page every agent publishes at a fixed URL. The seller looks that up in the AAMP registry, and the answer caps whatever the buyer claimed. An agent the registry doesn't know drops to the public floor, a blocked one is refused before any pricing goes back, and a buyer with no card and no API key can't be checked at all, so it floors to public too. Every check is recorded: what was claimed, what was enforced, which endpoint. One gap is left and the commit says so, because the shared negotiation message carries no field for the card address, so on that route the seller falls back to an API key.
A buyer agent used to be able to claim a better tier and get a cheaper price for it. Now the seller checks that claim against the registry, so the price matches who the buyer really is, and both sides can trust the number.
Bedrock Platform
View on LinkedIn
M
Magdalena Śleboda
Head of Operations | Scaling Tech Organizations with AI, Automation & Data-Driven Execution | Global Ops & Transformation Leader
3 weeks ago in 73% of our engineers who've been at Appliscale two years or more have spent those two years on the same client team.
For a consulting company, that isn't an HR metric. It's what the client is actually buying.
Because the thing that breaks a long engagement is turnover. The engineer who knew the system leaves. The context walks out with them. Three months of ramp-up, and someone pays for it.
So we screen against our own short-term interest. A candidate meets the people they'd actually work with, not only a recruiter. Our engineers ask the questions. They also get to say no.
That costs us. Turning down a strong engineer who isn't the right fit means a seat stays open, and an open seat is revenue we don't book this quarter.
We do it anyway. The alternative is more expensive, just later, and on the client's side.
Twenty of our engineers have been with the same client since at least 2023, and fourteen of them since 2022. On one account, the client changed its name along the way. The team didn't.
We automate and simplify every process we can. Not this one. No trade-offs here.
Retention is part of our culture and part of what we deliver.
View on LinkedIn
Damian Naglak
Head of Engineering | Bedrock Platform | AdTech
3 weeks ago in Embedding ten million pages costs ten dollars with a small model and about two hundred with a bigger one, on the same machine with the same pages.
On Tuesday I timed how long it takes to turn one page into a vector. Those timings turn into a bill, so here it is. I rented one entry-level graphics card at about 85 cents an hour and fed it pages in batches of sixteen, which is how you would run this if you were working through a catalogue rather than answering a live request. Ten million pages came to roughly $10 with the smallest model I tested, $30 and $45 with two mid-sized ones, and $220 with the largest. Small and large here means the number of values the model learned during training, from 33 million up to 600 million, and more of them generally means a better model and a slower one.
All four models are free to use commercially, so the machine rental is most of what you pay. Two other things move the bill as much as the model does. Sending sixteen pages through at once instead of one was worth up to 12 times, because the card sits mostly idle on a single page. Storing each number in the vector in less space is worth another two to three times, going by published results.
So is running your own cheaper than buying it in? Yes, as long as the card stays busy. Priced the same way, per million words of text, the biggest model I ran cost four and a half cents, against thirteen to fifteen for cloud providers. The smaller models came in under a cent. The exception is OpenAI's small model at two cents, which undercuts everything I ran.
The catch is that a rented card is billed by the hour whether it does anything or not, and those figures assume it never stops working. Start a job overnight, leave the machine up all day, and the card works under a fifth of the time, so the real cost is five or six times higher and the hosted price wins. Run the job hard and hand the machine back and the figures hold. For the biggest model the line sits near a third: busier than that and running it yourself is cheaper than Google's price, below it you are paying for an idle machine.
Where it starts to add up is refreshing user profiles rather than embedding a catalogue once. Re-embedding tens of millions of histories every day is a standing job rather than a one-off, so you are renting cards continuously, and the bill lands somewhere in the tens to hundreds of thousands a year depending on how many users you carry, how often you refresh, how long each history runs and which model you point at it. The direction that changes the picture is a bigger model. Cost tracked parameter count closely across the four I ran, so fine-tuning something several times larger multiplies all of this in step. Long histories push the same way, because past a few thousand words the cost climbs faster than the text does.
Bedrock Platform
View on LinkedIn