Financial Infrastructure · DevOps · Boston

Engineering inside
financial services.

Ten years across private equity, retail, and asset management. Real technical experience covering infrastructure, cloud, security, and trading systems. Written plainly to help other engineers navigate this world.

Michael Harlow
Michael Harlow // sys.ghost  ·  Boston, MA
☕ Buy me a coffee
Latest post
Every engineering org I talk to is running AI agents somewhere in production now. Almost none of them have figured out how to operate them the way we operate everything else that touches customer data.
Jul 14, 2026 · 11 min read
Read →
Latest post
Every engineering org I talk to is running AI agents somewhere in production now. Almost none of them have figured out how to operate them the way we operate everything else that touches customer data.
Jul 14, 2026 · 11 min read
Read →
All posts

Archive

We scheduled a Technology Connect event for our engineering team and built a lineup of games around AI and infrastructure topics. Here's how we put it together and what I'd do differently.
← Back to posts
Finance Tech May 1, 2026 · 11 min read

Why AI Vendors Are Moving to Token-Based Billing, and What It Actually Costs You

Why AI Vendors Are Moving to Token-Based Billing, and What It Actually Costs You

A few months ago I sat in a budget meeting where someone asked a simple question: what are we actually spending on AI tooling per user per month? Nobody in the room could answer it. Not approximately. Not even within an order of magnitude.

That is the moment I started paying closer attention to how enterprise AI billing is evolving, because the answer to that question is becoming genuinely difficult in a way it was not two years ago.

The old model and why it worked

For most of enterprise software history, the billing model was simple. You paid a per-seat license, annually, negotiated with procurement, predictable in the budget, boring in the best way. If you had 500 engineers and a tool cost $50 per seat per month, you paid $25,000 a month. Done. No surprises.

When AI tooling started entering the enterprise, most vendors defaulted to this model because that is what enterprise procurement knows how to handle. GitHub Copilot launched at a flat per-seat rate. Most AI-augmented SaaS tools bundled AI features into existing seat pricing. It felt familiar.

That model is quietly breaking down.

What is changing and why

The economics of serving language models do not fit the per-seat model well. Usage varies enormously between individuals. An engineer who uses an AI coding assistant lightly and one who runs large context windows all day do not cost the vendor remotely the same amount to serve. At small scale this variance averages out and the flat rate works well enough. At enterprise scale with thousands of seats, it stops averaging out and starts creating problems.

Vendors have been absorbing this mismatch, and some are deciding they can no longer afford to. The result is a gradual shift toward consumption-based billing: you pay per token generated or processed, and your bill reflects what you actually use.

This shift is happening unevenly. Some vendors are making it explicit. Anthropic, OpenAI, and Google have always billed API access by token for direct API users. The change is that the tooling built on top of those APIs is increasingly passing token costs through to enterprise customers rather than bundling them into flat seats. And vendors building native AI products are starting to move toward token or credit consumption models for their enterprise tiers.

The reasons are straightforward. Token-based billing aligns costs with usage, which benefits vendors whose cost structure is consumption-based. It also enables tiered pricing based on model capability: a query handled by a smaller, cheaper model costs less than one routed to a larger, more capable model. That granularity is hard to express in a per-seat structure.

What a token is and why it matters for budgeting

If you have not worked directly with language model APIs, the token concept is easy to underestimate.

A token is roughly equivalent to three or four characters of text, or about three quarters of a word in English. A short query like "summarise this email" is maybe ten tokens. A legal contract fed into a model as context might be fifty thousand tokens. A model generating a long response to that contract might produce another several thousand tokens. Input and output are typically priced separately, with output tokens more expensive.

For light, conversational use cases, token costs are genuinely negligible. For heavy enterprise use cases involving large context windows, document processing, or agentic workflows where models chain multiple calls together, the numbers can grow in ways that are not immediately intuitive.

The specific pattern that catches people off guard is agentic AI workflows. If you are building an automated process where an AI model breaks a task into subtasks, calls tools, reads results, synthesises information, and iterates, each step in that chain is generating and consuming tokens. A workflow that feels like one task might be twenty or thirty API calls with the full conversation context passed each time. The token count can be orders of magnitude higher than a simple query-response pattern.

The repercussions for finance and operations teams

In financial services, where I spend most of my time, a few specific consequences are worth thinking through carefully.

Budgeting becomes genuinely hard. Per-seat costs were forecastable. Token consumption depends on how employees use the tools, what use cases get adopted, how deeply AI gets embedded into workflows. Finance teams are used to modelling software costs as headcount times a rate. Token costs require modelling usage patterns that are often not well understood at the point of budget submission.

I have seen two failure modes here. The first is underestimation: a team gets token access approved at a low forecast and then a single high-usage team or a poorly bounded agentic workflow blows the budget in a quarter. The second is overestimation: risk-averse teams provision a token budget ten times what they actually use, the vendor sees low utilisation and questions the renewal, and everyone ends up in an uncomfortable conversation about value.

Cost attribution gets complicated. With per-seat billing, you know who is using what because you know whose seat is active. With token consumption, especially through shared integrations and automated workflows, tracing cost back to business units or projects requires deliberate instrumentation. If you are running a shared AI service layer that multiple teams call into, the billing comes in as a lump sum unless you have built cost attribution into your architecture.

Compliance and audit trails intersect with billing data. Regulated firms often need to demonstrate that AI-assisted decisions were made appropriately, with appropriate human review. In a token-based billing environment, the same logging infrastructure that supports audit trails can also support cost attribution. Getting these aligned from the start is much easier than retrofitting it.

Vendor negotiations shift. Enterprise software negotiations typically involve volume discounts tied to seat count. Token-based negotiations involve committed spend levels with discount tiers, minimum consumption commitments, and model-tier pricing that can be complex to evaluate without detailed usage data you may not yet have.

What is actually fair about this model

I want to be honest here: I think consumption billing is a more rational model than flat-seat licensing for AI tooling, even though it creates operational complexity.

Flat-seat pricing for AI tools subsidises heavy users with light users. In a large enterprise, that means teams doing serious, high-value AI work are partly paying for teams who have a license but barely use it. Consumption billing means you pay for what you use, which is a better signal for understanding where AI is actually delivering value.

It also unlocks access. A flat-seat model forces you to decide upfront who gets access. A consumption model lets you give broad access and observe where usage concentrates, which is often more valuable for understanding adoption than any survey or internal roadmap exercise.

The problem is not with the model itself. The problem is that most enterprise teams are not ready for it operationally, and the tooling vendors have not done enough to help.

What you actually need to do about it

If your organisation is adopting AI tooling at scale, a few things are worth getting right before consumption billing catches you off guard.

Instrument everything from day one. Every AI service call should be tagged with at minimum: the team or business unit, the use case or workflow, the model version used, and the timestamp. This is not onerous if you build a thin shared client library that every team uses to make AI calls. It becomes very onerous if you try to retrofit it after the fact when your first large bill arrives.

Model your heavy use cases explicitly. Identify the workflows in your firm that involve large context windows, agentic patterns, or high-frequency calls. Estimate token consumption for those specifically. Light conversational use cases will not drive your bill; the edge cases will.

Negotiate committed spend, not seat count. When talking to AI vendors, push the conversation toward committed token or credit volume with defined pricing tiers rather than seat-based licensing. You get better rates on committed spend, and you have a ceiling to budget against.

Build in budget alerts. Both the major cloud providers and the AI API providers offer spend alerts. Set them at levels that trigger human review before they trigger a surprise at month end.

Consider a token budget layer for internal teams. If you are running a shared AI service, implement per-team token quotas with monitoring dashboards. This is the same problem as cloud cost management and the same solution works: give teams visibility into their consumption, set soft limits with alerts, and hard limits only where strictly necessary.

The longer-term picture

The move to token-based billing is one of several structural changes happening simultaneously in enterprise AI pricing. Model capability tiers are becoming more granular. Caching of repeated context is starting to appear as a pricing lever. Specialised models for specific domains are being priced differently from general-purpose models.

The net effect is that AI cost management is becoming a discipline in its own right, the way cloud cost management became a discipline around 2015 to 2018 as cloud adoption scaled. The organisations that built internal competency in cloud economics early ended up with meaningfully better outcomes than those who treated cloud billing as a finance problem rather than an engineering problem.

The pattern will repeat with AI. The engineering teams that understand token economics, instrument their usage, and architect their workflows with cost efficiency in mind will have a significant advantage over those that discover the problem through an unexpected bill.

I am building that instrumentation layer at my firm right now. It is not glamorous work. But I have sat in enough of those budget meetings to know that having the data when the question gets asked is worth considerably more than explaining afterward why you did not.

Found this useful?
☕ Buy Michael a Coffee
← More posts

Hey, I'm Michael Harlow.

Senior Systems Engineer · Boston, MA · Writing as sys.ghost

I have spent over a decade building and maintaining infrastructure at the intersection of technology and financial services. My career has taken me through three distinct sectors -- technology, private equity, and asset management -- and each one changed how I think about what reliable infrastructure actually requires.

I started in general IT, which is where most engineers who did not go straight into software end up. Data centers, networking, on-call rotations, learning to label cables properly because unlabeled cables are a promise that someone else will suffer later. The work taught me that almost every sophisticated system is, one layer down, a collection of unglamorous fundamentals that either hold or do not. I still believe that. I still label everything.

Private equity came next, and it was a different world. The infrastructure stakes there are less about uptime and more about data integrity. When deal teams are making acquisition decisions based on data you are responsible for, and when a due diligence process has a hard deadline that does not move regardless of what broke overnight, your relationship with reliability changes. A wrong number in an LP report does not cause an immediate incident. It causes a conversation in a partner meeting six weeks later, and by then you need to reconstruct what happened from imperfect records. I became obsessive about data provenance in PE and I have not stopped.

For the past several years I have been in asset management, supporting trading and investment operations infrastructure. This is the environment I find most technically interesting. The compliance requirements are demanding, the legacy systems have long institutional memories, and the tolerance for operational errors is genuinely low -- not just in terms of business impact, but in terms of regulatory consequence. When markets are open, there is no fixing it after the weekend.

I started Packet & Profit in January 2026 because I kept looking for the kind of writing I wanted to read and finding it mostly did not exist. There is a lot of content for engineers online. There is much less written by engineers working specifically inside regulated financial services firms, being honest about what that actually involves day to day. The compliance conversations, the legacy constraints, the incident management in front of stakeholders who measure downtime in dollars per minute. That is what I write about here.

Outside of work I have been running a Saturday morning robotics course at my local YMCA for kids aged 10 to 14. It is one of the better decisions I have made.

Certifications

Red Hat Certified Engineer (RHCE)
Certified Kubernetes Administrator (CKA)
AWS Solutions Architect -- Associate
CompTIA Security+
HashiCorp Vault Associate

My Stack

RHEL / Ubuntu
Kubernetes
OpenShift
Terraform
Ansible
Prometheus
Grafana
Python / Bash
AWS / Azure
Cisco / Palo Alto
PostgreSQL
Redis
HashiCorp Vault
Fluent Bit
Helm
ArgoCD

Career

2022 -- Present
Senior Systems Engineer, Asset Management -- Boston, MA
Leading infrastructure for trading operations and investment management systems. Responsibilities span network security, cloud migration strategy, Kubernetes platform engineering, and incident response. Deeply involved in T+1 settlement infrastructure work and the shift from overnight batch processing to near-real-time event-driven architecture.
2018 -- 2022
Systems Engineer, Private Equity -- Boston, MA
Built and maintained data infrastructure supporting deal teams, portfolio monitoring, and investor reporting. Managed infrastructure through multiple due diligence cycles with hard deadlines and high data integrity requirements. Led a major data platform migration from on-premises to cloud-hosted infrastructure, including security controls satisfying LP and regulatory requirements.
2015 -- 2018
Infrastructure Engineer, Retail Technology
Supported inventory management, real-time pricing, and supply chain integration systems across a high-SKU retail environment. Operated under peak load conditions where scale was a concrete engineering problem rather than an abstract one. Built out monitoring and alerting infrastructure from scratch and managed a full data center relocation.
2013 -- 2015
IT Engineer, Technology Sector
Established the professional fundamentals: data center operations, network infrastructure, endpoint management, and the on-call rotations that teach you more about system fragility than any textbook. Developed an appreciation for cable labeling that has never left me.

Get in Touch

If you are an engineer working in financial services, curious about the career path, or have a question about something I have written, I would genuinely like to hear from you. Use the and I will get back to you. If something here has been useful, a coffee is always appreciated.

A note on anonymity: I write under my own name but keep my current employer private. The financial services industry is small, the regulatory environment is real, and I want to write honestly without those constraints. All incidents and case studies on this site are anonymised. The technical content is real; identifying details are not.
Get in touch

Contact

Whether you are an engineer in financial services, have a question about something I have written, or just want to say hello - feel free to reach out. I read everything.

Powered by Resend · No spam, ever

Legal

Privacy Policy

Last updated: April 2026

This policy explains what information Packet & Profit collects when you visit this site, how it is used, and what choices you have.

Information We Collect

We do not require you to create an account or provide personal information to read this blog. The only personal information we collect is what you voluntarily submit through the contact form: your name, email address, and message. This information is transmitted via Resend and used solely to respond to your enquiry.

Google AdSense and Advertising

This site uses Google AdSense to display advertisements. Google AdSense uses cookies and similar tracking technologies to serve ads based on your prior visits to this and other websites. This means Google may use information about your visits to this site to show you personalised ads on other sites across the web.

You can opt out of personalised advertising by visiting Google Ads Settings, aboutads.info, or optout.networkadvertising.org. See Google advertising policies for more.

Cookies

This site uses a single first-party cookie to remember your theme preference (light or dark mode). This cookie contains no personal information. Third-party cookies may be set by Google AdSense for advertising purposes as described above.

Analytics

This site does not currently use any analytics platform beyond what Vercel provides as part of its standard hosting service (aggregated, anonymised traffic data).

Contact Form

When you submit the contact form, your name, email address, subject, and message are transmitted to the blog author via Resend. This data is not stored by this site and is not shared with any third party beyond Resend. See Resend's privacy policy for details.

Third-Party Links

Posts on this site may link to external websites. We are not responsible for the privacy practices or content of those sites.

Your Rights

If you have submitted a message via the contact form and would like that information removed, or if you have any questions about this policy, please use the contact form to get in touch.

Changes to This Policy

We may update this policy from time to time. The date at the top of this page reflects when it was last revised.

Legal

Terms of Service

Last updated: April 2026

By accessing and using Packet & Profit (packetandprofit.com), you agree to be bound by these Terms of Service. If you do not agree, please do not use this site.

Use of Content

All written content, illustrations, and code examples published on this site are the original work of Michael Harlow unless otherwise stated. You are welcome to share links to posts and quote brief excerpts (with attribution), but you may not reproduce full articles, copy content to other websites, or use the content for commercial purposes without written permission.

No Professional Advice

Content published on this site reflects personal opinions and professional experience. It is provided for informational and educational purposes only. Nothing on this site constitutes financial, investment, legal, or professional advice of any kind. See the for more detail.

Third-Party Links

This site may contain links to third-party websites. These links are provided for convenience and do not constitute an endorsement of the linked site or its content. We have no control over and accept no responsibility for external sites.

Advertising

This site participates in Google AdSense, which displays advertisements from third-party advertisers. The presence of an advertisement does not constitute an endorsement of the advertiser's products or services. Ad content is determined by Google based on the content of this site and your browsing history.

Accuracy of Information

While we make every effort to ensure the accuracy of information published on this site, technology and financial markets change rapidly. Information that was accurate at the time of publication may become outdated. We do not warrant the completeness, accuracy, or timeliness of any content on this site.

Limitation of Liability

To the fullest extent permitted by law, Packet & Profit and its author shall not be liable for any direct, indirect, incidental, or consequential damages arising from your use of, or inability to use, this site or its content.

Changes to These Terms

We reserve the right to update these terms at any time. Continued use of the site following any changes constitutes your acceptance of the revised terms. The date at the top of this page reflects the most recent revision.

Contact

If you have questions about these terms, please use the .

Legal

Disclaimer

Last updated: April 2026

Packet & Profit is a personal blog written by Michael Harlow, a Systems Engineer based in Boston, MA. The views expressed here are entirely his own and do not represent those of any employer, client, or organisation he is affiliated with.

Not Financial or Investment Advice

This site discusses financial services technology, investment management infrastructure, and related engineering topics from a technical practitioner's perspective. Nothing published here is financial advice, investment advice, or a recommendation to buy, sell, or hold any security, asset, or financial instrument. The author is not a registered financial adviser, broker, or investment professional.

Content that references financial markets, trading systems, or investment firms is provided for technical and educational context only. Any figures, case studies, or examples are illustrative and should not be relied upon for financial decisions.

Not Legal or Professional Advice

Nothing on this site constitutes legal, compliance, regulatory, or professional advice. Readers should consult qualified professionals for advice specific to their circumstances.

Professional Experience

Posts on this site draw on the author's professional experience in systems engineering across private equity, retail technology, and asset management. Specific details about employers, clients, projects, and colleagues have been anonymised or generalised. Any resemblance to specific organisations is incidental.

Accuracy

The author makes reasonable efforts to ensure published information is accurate at the time of writing. The technology and financial services landscape changes quickly. Readers should verify any technical or regulatory information against current primary sources before acting on it.

Affiliate Links and Advertising

This site displays advertisements through Google AdSense. The site may also contain links to tools, services, or products that the author uses or finds useful. These are not paid endorsements unless explicitly stated. The author's opinions are his own and are not influenced by advertisers.

Questions

For questions about anything on this site, please use the .

This site uses cookies for theme preferences and displays ads via Google AdSense, which may use cookies to personalise ads.