If we need to top up our AI account every few days, it makes sense to look for a cheaper model. Especially when much of the work repeats: reading files with similar formats, extracting the same information, and producing reports with the same structure.
One argument is that a cheaper model should be enough if we are good at writing instructions and skills. I think this is worth examining, because preparing those instructions takes time too. Someone has to find examples, try them, correct what went wrong, and run the work again.
So we have two costs to consider: what we pay to use the AI, and how much work we do to get a result we can use.
Let's follow one type of task through that calculation. Every amount and duration in the example is an assumption for illustration. These are not prices for a particular model or operating figures from a company.
What Can a Skill Help With?
Suppose our team receives weekly sales files from several branches. We want AI to read them, calculate net sales, and explain which branches performed very differently from the previous week.
The instructions need to cover a few specific details. Are cancelled transactions still in the file? Have discounts already been deducted? Should differently written branch names be combined? Does a refund belong to the week of the original sale or the week the money was returned?
Someone who regularly prepares these reports probably knows the answers. We can package those rules, sample files, and scripts for repeated calculations into a skill that the AI can use.
A skill here is a package of instructions and supporting resources for a particular task. Creating one does not automatically change the underlying model through training.*
The AI then has less to work out from scratch each time a file arrives. To calculate totals, for example, it can run a script we have checked and use the output to write its explanation. A cheaper model might handle this well.
SkillsBench compares agents with and without prepared skills. The research finds that skills can improve task completion, and that some smaller models with skills can match larger models without them. The benefit varies across tasks, though. Asking models to generate their own skills does not automatically improve performance. SkillsBench paper.
That gives us a reason to test a cheaper model with suitable skills. For our sales reports, we still need to check the totals, the treatment of refunds, and the explanation. A benchmark score alone does not tell us how long our team will spend getting that result.
What Does Preparation Cost?
Suppose we test two options on the same reports. Option A uses a more expensive model. Option B uses a cheaper model with additional skills and dedicated scripts.
Both need business rules, data access, and checks on the result. To keep the comparison manageable, we will leave out costs that are identical in both options and count the costs that differ.
Assume A takes 4 hours to prepare and test. B takes 20 hours because we need additional scripts, examples, and instruction changes to meet the same requirements.
At an assumed cost of Rp 150,000 per hour, preparation for A costs 4 times Rp 150,000 = Rp 600,000. Preparation for B costs 20 times Rp 150,000 = Rp 3 million.
B requires Rp 2.4 million more in preparation. If we do the work ourselves, that amount does not arrive as a separate invoice. But the extra 16 hours are still used up, and could have gone toward other work.
Those hours can differ considerably. The expensive model may need extensive corrections too, or the skill we need may already exist. The 4 and 20 hours are assumptions for this example. In our own work, we should record the time spent preparing both options.
How Often Must We Use It Before It Costs Less?

After testing, suppose the AI usage cost to complete one report is Rp 5,000 for A and Rp 1,000 for B. These amounts include all model calls and retries, assuming the average stays constant over the comparison period.
For now, assume the review time and final quality are the same. B saves Rp 4,000 per report.
Divide the extra Rp 2.4 million in preparation by Rp 4,000 in savings, and we need 600 reports to recover the difference. At report 600, total costs are equal. B becomes cheaper after that, as long as our assumptions hold.
At only 100 reports, A costs Rp 600,000 plus Rp 500,000 = Rp 1.1 million. B costs Rp 3 million plus Rp 100,000 = Rp 3.1 million. B's AI usage is 80% cheaper, but its total cost is still Rp 2 million higher.
At 1,000 reports, A costs Rp 5.6 million and B costs Rp 4 million. By then, the usage savings have paid for the additional preparation.
This is the break-even point: the volume at which both options have the same total cost. If we need 20 reports per month, reaching 600 reports takes 30 months. At 1,000 reports per month, we reach that volume within the first month.
The 30-month calculation assumes we can keep using the skill that long without extra costs from changes. File formats, business rules, or the model itself may change in the meantime. How often we use the skill matters to whether the preparation is worth doing.
What If Review Takes Longer?

Let's change one assumption. B needs an extra 2 minutes of review per report because the team more often has to check its explanation and correct figures quoted incorrectly from the script output.
At Rp 150,000 per hour, one minute costs Rp 2,500. An extra 2 minutes costs Rp 5,000 per report. The AI usage saving was only Rp 4,000, so B now costs Rp 1,000 more for every report, even before we count its additional preparation.
If that continues, higher volume will not make B more economical. Every additional report increases the difference.
But we need to measure those 2 minutes. We cannot assume that a cheaper model always makes more mistakes. B's scripts might produce more consistent calculations and reduce review time. A can also quote the wrong figure or miss a refund.
We can compare them using the same set of reports and the same acceptance criteria. All transactions must be accounted for, net sales must agree with a checked calculation, and the explanation must refer to the correct branches and periods.
Record the cost and time until the report meets those requirements. Failed attempts still count, including cases where someone finishes the report manually. If we request 100 reports and only 80 are completed, we must also account for the remaining 20. An option should not appear cheaper because unfinished work was removed from the calculation.
There is also a difference between the value of time and cash savings. If the team's salaries stay the same, reducing review time may not lower spending that month. It may give people time for other reports, reduce overtime, or delay the need to hire another person.
Can We Keep Using the Skill?
Return to the original example, where review time is the same. Suppose B needs 4 additional hours of maintenance per month compared with A. Perhaps a branch changes its file format and we need to adjust the script.
That costs 4 times Rp 150,000 = Rp 600,000 per month. At a saving of Rp 4,000 per report, we need 150 reports per month just to cover the additional maintenance.
At 100 reports per month, AI usage saves Rp 400,000, while maintenance costs Rp 600,000. B costs Rp 200,000 more each month, and we have not recovered the extra Rp 2.4 million in preparation either.
At 1,000 reports per month, AI usage saves Rp 4 million. Subtract Rp 600,000 for maintenance, and the monthly saving is Rp 3.4 million. That covers the extra Rp 2.4 million in preparation within the first month, including a full month's maintenance cost.
Some skills can serve several teams, spreading their preparation cost across more work. Those tasks should use the same rules and scripts, though. If each team needs its own adjustments, we need to count that time too.
Which Model Should We Start With?
I would start by trying both options on work that represents what the team handles regularly, including files that have caused problems before. That helps us see which parts need a more capable model and which benefit from clear instructions or scripts.
The skill built for B may help A too. We can then compare both models with that skill. There is no reason to keep the expensive model without skills just because that was our starting comparison.
For work we will do only a few times, preparing and maintaining a skill may cost more than it saves in usage. Paying for a more expensive model can make sense if it gets us an acceptable result sooner at a lower total cost.
For repeated work with reasonably stable rules, there are more opportunities to reuse the preparation. A cheaper model with skills becomes an attractive option if the results meet our requirements and the total cost, including review, remains lower.
We can also use both in one process. A cheaper model handles familiar file formats, while files that fail checks go to another model or a person. The checks and escalations have their own costs. We also need to look for errors that pass those checks, since a more expensive model will not necessarily correct them on its own.
To begin, we could take 30 old reports, use some to improve the skill, and reserve the rest for a final test. Record preparation time, AI cost, review time, and the number completed for both options. This is an initial way to find recurring problems, not enough evidence that rare failures are covered. Once we have those figures, we can compare them with the number of reports we will actually need next month.
