Payments & cashier
Deposits and withdrawals by card, e-wallet, crypto and bank transfer. 3DS checks and drop-offs, retries after a decline, double clicks and double charges, currency rounding, fees, limits and reversal windows.
Tbilisi, Georgia | Remote | UTC+4
I work on the flows where money and licences live: sign-up, KYC, deposits, first deposit tracking, game settlement, bonuses and withdrawals. Seven years running an online casino back office, then QA on the same flows. Fourteen years in operations and QA, eight of them in regulated casino and sportsbook platforms.
Nothing here is a mock-up. Every number below comes from a live platform I build, run and test.
Background
Every flow on this page, I ran in a live casino before I tested it.
For seven years I worked in the back office at Green Feather Online, a licensed casino operator: five as a shift manager, the last three as the only one, running a desk of up to eleven agents while still handling cases myself.
The desk owned every withdrawal from the request until it was approved or refused; finance then sent the money. The decision was ours, and it had to rest on evidence.
Finding that evidence is most of the job: account activity, login history, linked accounts and game history. Any one of them could turn a routine payout into a case worth stopping.
Documents were the other half, and the rule was not obvious. What we asked for depended on the deposit and payment method, not on KYC status. A withdrawal in EUR or USD needed far more evidence than a payout to a local-currency account, so the same player could pass for one route and fail for another.
When payments broke, I owned the case from start to finish, taking it to finance or straight to the payment provider. I tracked which methods and currencies actually worked in each market, in a region where few providers could pay out at all, and answered the affected players myself, by email.
Some numbers from that desk, mine rather than the team's. In a typical year I caught five to ten forged identity documents, and stopped stolen-card deposits at least twice, including attempts using three to five stolen cards on one account. Three to five times a year I stopped a withdrawal where a player had used a bug in a game to build the balance. And I found a software bug that was quietly overpaying customers, forty to fifty euros each.
That last one is why the rest of this page exists. It was a QA finding, made years before I had the job title, from an operations seat, and nobody had asked me to look.
Fraud was constant, not occasional. Before all of that, six years in retail banking, where every entry can be audited and a mistake is a regulatory event, not a bug.
That is why the testing on this page focuses where it does. I am not guessing which parts of a payment journey fail. I spent seven years answering for them.
What I do
In iGaming a bug is rarely just a bug. It is a chargeback, a stuck withdrawal, an abused bonus or a compliance problem. I test for that.
Deposits and withdrawals by card, e-wallet, crypto and bank transfer. 3DS checks and drop-offs, retries after a decline, double clicks and double charges, currency rounding, fees, limits and reversal windows.
Document upload and rejection loops, limits before and after verification, repeat KYC triggers, age checks by country, self-exclusion checks, and the withdrawal gates that depend on all of it.
Provider games loading, real versus demo mode, one bet charged once, round settlement and balance sync, lost connections mid-round, and the network switches that cause double charges.
Contribution rates by game type, which balance is spent first, max-bet breaches during wagering, expiry and forfeit rules, and the abuse paths that turn a promotion into a loss.
Checks across browsers and real devices: iOS Safari and app quirks, private-browsing failures, rotation and screen-size traps, slow networks, and tap targets big enough to hit.
Edge-case testing at the API level, regression suites proven to fail before they are trusted, and end-to-end tests that run in a separate environment, never on production data.
Method
The path every player and every euro takes. Pick a stage to see what I check, and what goes wrong there.
First impression, and the first place fraud gets in.
Where withdrawals go to die when it is built wrong.
The most valuable flow on the platform.
Where marketing money is either earned or quietly wasted.
Someone else's games, your liability.
The most abused part of any casino platform.
Every earlier bug shows up here, at the worst moment.
Not a feature. A licence condition.
Proof of work
Most QA portfolios show made-up sample data. This one does not: Michi is a job-matching platform I built and run, and every number below is measured from its code.
It reads a CV and returns a ranked list of jobs the person can realistically get, with the reasons shown. Python and FastAPI on the backend, Next.js and TypeScript on the front, PostgreSQL, scheduled imports and a human review queue. michi.ge
27 bugs, each with root cause, fix, and clear limits. All closed.
Steady delivery, not a weekend project.
In June I audited my own test suite. It rated the front ends as untested. This is that gap closing.
| Surface | At audit (June) | Now (August) |
|---|---|---|
| Backend | 85 test files | 171 files / 4,862 tests |
| Candidate frontend | ~5 files | 63 test files |
| Admin console | 0 behavioural tests | 45 E2E tests |
On disclosure. Michi's code is private. Everything here was checked before publishing, so no scoring weights, thresholds or internal details appear on this page. I treat client information the same way.
Case studies
Five from BetMavrik's platforms, four from the job-matching platform I build and run. Written from my own records, with product internals and third-party names removed: how the bug was found, how the cause was proven, and how the fix was checked.
Reported to the vendor, their fix checked live, then locked in by an automated test
On BetMavrik's lottery product, the game API accepted bets on a draw that had already closed. That is the one thing a lottery cannot get wrong: once the entry window shuts, no more money may be taken.
A bet taken after close is a money problem, not a display bug: the operator pays out on a result someone may already know, or faces a complaint that reaches the regulator. And it is silent. No alert, no error, and the product looks healthy until someone checks the tickets.
The API belonged to a vendor, so I could not fix it. That changes the job: describe it clearly enough for someone else to act on, then check their fix instead of trusting it. I wrote a full incident report, and confirmed the fix live in a separate retest.
I then turned the case into an automated test, so the build fails if the bug comes back. It runs in the product's daily checks. A vendor fix is only proven for the day you checked it: their code can change without telling you, and if this bug returns, it looks like normal traffic.
The rule I work by. Any bug fixed by someone else gets a test on your side, or it is not closed. A promise is not a control, and "we fixed it" has no expiry date.
I checked the automation against its own screenshots, then checked my own fix and found it did not work
A crawler opened every game in a casino's catalogue and reported whether it loaded. But each game runs inside a window served by another company, which our code cannot see into or click. So all the crawler could really check was that the window existed. The screenshot it saved was the only real evidence, and nothing was reading it.
I compared its verdicts against its own screenshots:
| Site | Self-reported | Checked by hand |
|---|---|---|
| First | ~88% passing | ~36% |
| Second | ~74% passing | ~48.5% confirmed working |
Two causes, and the window check could not see either. Some games showed a loading screen that never finished, with the window there the whole time. Others had loaded, but sat behind a start screen the crawler could not click past.
I added an automatic click for the second cause. Checking every one of roughly 2,300 screenshots proved it runs on every game and still does not work: it clicks the centre of the window, while the real "continue" buttons sit near the bottom. Right coverage, wrong pixel.
The stuck loading screen is still unsolved. I built a detector that reads the text on each screenshot, tested it on all of them, and rejected it because it missed too many working games. Both gaps are written down as open.
The crawl runs on a test account with no money, so games that loaded fine then showed a "no funds" message over working reels. My first check read that message as an error and failed about 110 healthy games. The crawler had been right; I was wrong. Checking by hand can be wrong too.
Ask this of your own suite. A pass rate is a claim your automation makes about itself, and almost nobody checks it. If it saves screenshots or logs, you already have the evidence to test whether its greens are real.
Four ways the results page lied, in a dashboard I built myself
I built the dashboard that runs our browser test suites: a card per suite, click to run, and a coloured result at the end. Every bug below is one I shipped and then found in my own tool.
1. The badge and its own numbers judged the same run by different rules. The badge came from the test runner's result, which ignores the health-check and balance rows. The numbers underneath counted those rows as failures. So a card could show a red "1 failed" line right below a green "Passed" badge.
2. A run where nothing happened looked like a clean pass. With zero tests counted, the card showed zero of zero, in green.
3. One run could silently overwrite another's report. The daily and regression cards shared a folder and wrote to the same default file. A run started from a terminal wrote straight over the daily report.
4. The results page hid two of its own categories. It showed Total, Passed, Failed and Skipped, but there are six categories, so Warning and Flaky were invisible even while the card beside them said "1 warn". Its label read "98% passing (54 / 56)", but 54 of 56 is 96%. The percentage and the label counted different totals.
Bugs 2 and 3 combined into a real incident. An interrupted run, 34 tests, none of them run, overwrote the daily report. The daily card then showed zero of zero passed, in green, under a badge left over from an earlier run that really had passed. No error, no alert: a healthy daily check that had never run.
The rule: the badge comes from the run's own result, and the numbers must agree with it. So every row goes into one of four groups instead of two:
| Group | What lands there | Counts toward pass rate? |
|---|---|---|
| passed / failed | Real checks: login, each section | Yes |
| warning | Health probe, balance detection | No, information only |
| skipped | Known failures, known-bug trackers | No |
| flaky | Failed, then passed on retry | Own group, never counted as passed |
The percentage now uses passed plus failed only, so a clean run reads 100% and matches its green badge. Each type of run writes to its own folder, so no run can overwrite another, including one started by hand.
The health check is a simple request that connects differently from the browser running the test; when it reports an error, that is often a bot check the real browser passes. The browser decides whether a page works, not the probe. So a probe that disagrees is information, not a verdict. Counting it as a failure would teach everyone to ignore red, which is as bad as a green that means nothing.
Check the thing you check with. Nobody tests the reporting, because it is the tool rather than the subject. If your dashboard shows "nothing ran" the same as "everything passed", every green you acted on was worth less than you thought. This one is not finished either; the remaining gaps are written down rather than quietly closed.
Six secrets needed, none of them set, and nothing showing the failures
A scheduled job ran the smoke tests against our live sites every day, saving reports and screenshots. It had been failing for at least seven days in a row, and nothing showed it.
The job needs six stored secrets: five logins and one notification channel. None of them existed. So every run failed at sign-in, every day, for a week. The tests themselves were fine; the job turned red only because there were no logins to sign in with.
"The tests failed" and "the tests never ran" produced the same red. They are not the same event and they do not have the same fix. One is a bug in the product. The other is a week with no checks at all on sites that handle money. And a red nobody reads is not a warning at all.
Only the repository owner can add these values, so there was no code to write. The result was a clear request to the one person who could act: which six secrets, what each is for, and what stays broken until they exist.
An earlier note in the project said only two of the six were missing. The real number was six. The note had gone stale, and I had to re-check and correct it before I could trust my own report. That correction is why the finding is right: written facts go stale, and the worst ones are those written by someone people trust, because nobody re-checks them.
This is Medium on purpose. Nothing in the live product was broken; what was lost was a week of checking. I would rather rate it correctly than inflate it. If everything is critical, nobody can sort the list, the same failure as a dashboard where everything is green.
Two questions for your own CI. When it goes red, can it tell "failed" from "never started"? And who receives that red, by name? A scheduled job with nobody watching its failures will one day spend a week failing where nobody sees it.
On a suite that spends real money every run, four tests that could not fail
Most test suites cost you time. This one spends real money: every run places real bets from a funded account. A test that checks nothing is normally wasted time; here it also spends money while reporting success.
A test named "should show an error when the stake exceeds the balance" could never show it. The game has a fixed price and no stake field, so that path could never run. Instead it placed an ordinary charged bet on every run, and one path logged "PASS" without checking anything. It had a green result, a believable name, a real cost, and no value. I rewrote it to call the game API directly and check that an over-balance stake is refused with no charge.
The new test carries a warning about its own worst case: if the backend has no balance check, this drains the wallet rather than costing one bet. A test that probes a money limit can cost more than the bug it finds.
A check that could never fail. The duplicate-bet test counted success messages with a method that only ever returns zero or one, then checked the count was never two. That can never be false. So duplicate-bet protection, the thing between a player and a double charge, was untested.
And it accepted almost any charge. The same test passed on any deduction below a generous limit, so two real bets went through happily. I rewrote it to count tickets before and after, and to require exactly one.
One endpoint returns the balance in hundredths of a unit; another in whole units. A check subtracted one from the other, comparing different units, so it could pass or fail for the wrong reason. Fixed by reading both the same way. This was a bug in the test, not the product, and I record it on purpose.
The script that ran every suite included the bulk ticket tests. So a casual "run everything", the most harmless-looking action there was, spent real money. It is now limited to the daily set and cannot spend.
Set-up created five tickets per game type, across three types. Cutting that to two roughly halved the real-money cost of a regression run. The catch: the suite's "all draws closed" check used a limit tuned to the old count. Left alone, a genuinely closed draw would have turned from a clean skip into a red failure. Adjusting that limit was what stopped the saving breaking the reporting.
Two questions for any suite that touches money. Can each test actually reach the state it claims to check, or is that path blocked and the test quietly doing something else? What does one full run cost, and who last asked?
A list with no fixed order feeding a scorer where position matters, found by comparing snapshots
I keep dated copies of all results, so any change can be compared before and after. After a routine change I did the usual check: explain every difference, or do not ship. Most explained themselves. Then 81 differences remained with no change in the inputs, 67 of them score moves on records nobody had touched. Some results appeared from nothing; others vanished. The clue: 35 of the moves were the same size, to one decimal place, in both directions.
Before hunting for a cause I ruled out the obvious ones. Every profile was exactly the same. The admin log showed no changes to any affected record. All job fields were identical. Three checks, one conclusion: something was changing the output while every input stayed still.
The scoring code built a list using a Python set, then read it back in order. The scorer cares about position: the same item is worth more in an earlier slot. But a Python set has no fixed order. It depends on a start-up setting that was not fixed anywhere, so every restart reshuffled it, and the nightly job wrote that shuffle into the database.
Two more effects made it worse than cosmetic. Records near a cut-off crossed it, so results really did appear and disappear for users between deploys. And it added constant noise to my own before-and-after comparison, the tool I use to approve changes.
A likely cause is not a proven one. I rebuilt it from clean test data and ran one pair of records under 25 different start-up settings:
| Condition | Result across 25 settings |
|---|---|
| Before fix | 27.7-point spread, in even steps |
| After fix | One identical score |
Read the list in sorted order, in three places, not one. The third place was found by searching the code's structure rather than its text: it saved an ordered list to the database, rewriting that field on every re-score with no real change behind it. A plain text search would have missed two of the three.
A normal test cannot catch this: the setting is fixed when the program starts, so every test in one run shares it. The regression test starts several separate runs with different settings and requires identical output. Two rules: it was proven to fail on the broken code first, and it ships with a second test confirming the example still depends on order, so the check cannot quietly stop meaning anything.
The measure of success mattered: two full re-scores in a row must produce identical stored results. Reading a saved copy cannot prove that. The overnight run moved exactly what the approved change predicted, same count and same pattern, so it added nothing of its own. And the clue was gone: not one difference at the tell-tale size, which had produced 97 the night before.
Why this matters for iGaming. Bugs that behave differently each run are the most expensive to find late, because they look "flaky" and get ignored. If anything with no fixed order feeds something where order matters, such as odds, bonus order or settlement order, you have this bug and do not know it.
A rule that turned into a coin flip, and still sounded certain
One listing was dead but still being shown. The expiry check had not caught it, and it was not honestly saying "I cannot tell". It had run, decided, and got the wrong answer. It saw a page that returned success, redirected to the employer's main page, carried an error flag in the address, and had a title that was not the job. Four signs saying "gone". Verdict: keep.
The rule was: if a strong job identifier survives into the new address, the listing has
only moved. Sensible. But the check for "is this part an identifier?" said yes if it
contained a digit or was 12 or more characters long, a length rule meant for
sites that use words as identifiers. These addresses look like
/{employer}/jobs/{id}, and employer names are usually long
too. So the rule saw an identifier survive and decided the job had moved. What
had really survived was the company name, which always survives, because that is exactly
where a site sends you when a job is gone.
Same redirect, same everything, changing only the employer name in the address:
| Employer name length | Dead job detected? |
|---|---|
| 6 characters | Yes |
| 6 characters | Yes |
| 17 characters | No |
| 19 characters | No |
That one table turned a vague "this seems unreliable" into a fix everyone agreed on in about thirty seconds. Then I measured how far it spread rather than guessing: roughly one in five affected listings had a name long enough to hide the problem.
Require the most specific identifier to survive: a part with digits in it
is the real listing id, so require that and ignore the words. Keep the length
rule only where no numeric part exists. And read the ?error flag, which had
been thrown away because only the URL path was compared.
This change turns a "keep" into an "expire", and that direction is dangerous. A wrong expiry silently deletes real listings, and the original caution was deliberate. So it did not ship on a passing unit test, but after a replay against the real data:
| Replay outcome | Count |
|---|---|
| Newly expired (each checked dead by hand) | 3 |
| Wrongly relaxed | 0 |
| Side effects | 0 |
Two lessons that transfer. A rule built from "any of these signs" becomes a coin flip as soon as one sign becomes common noise, and it does so quietly, still sounding confident. And when a fix makes a system less cautious, unit tests are not enough: replay it against real data and count what else changed before shipping.
"Could not check" and "checked, all fine" shared one path. 68% dead, reported green
A user reported a link opening a site's home page instead of the job. A small thing, and I had seen one similar report days earlier. Two stories are not a bug report, so I measured: on that source, 68% of live listings were dead. I confirmed it three ways before believing my own number: a bulk check, a one-by-one recheck, and a real browser, in case the bug was in my own tool.
The check expired a listing on a few specific "not found" responses. Every other outcome fell through to "keep". That one default merged two completely different states: confirmed working and could not check.
The source answered with a bot-check page instead. That was neither "not found" nor success, so the first branch missed it and the second never ran. Verdict: keep. Forever. A bot-blocked source can never be checked, and its dead listings pile up unseen: no counter, no log line anyone read, no health check. Looking for others with the same shape, I found a second. Together, a quarter of all live listings sat in sources that had never once been checked successfully.
The obvious fix, detect the bot check and retry, does not work, and proving that took a real experiment. The bot check appeared for dead and live listings alike from the server, while the same addresses worked normally from a home connection. So the block was on the server's network, not on specific pages. Pretending to be a normal browser was ruled out by test, not assumption: a real browser identity was blocked in exactly the same way.
Labelling the response "could not check" makes the blind spot visible, but does not bring back the ability to check. Those were two separate problems, and I stopped treating them as one. So: fix the data at once and make it reversible, give the check three outcomes instead of two, and plan the different connection as its own piece of work. One rule written down: never expire automatically on "could not check". Being too strict there deletes real listings, and you never find out.
Go and read your monitors. Every health check you own has a default branch. Ask what it does when the check could not run, not when the check fails. If "I could not reach it" and "it is fine" produce the same green, that monitor will one day reassure you straight through an outage. That is worse than no monitor, because a missing monitor does not lie to you.
A hard question, an honest answer, and the gap closed. 86 audits so far
Michi runs an ongoing audit programme: 86 written audits covering more than 60 areas, including accessibility, security, performance, GDPR, reliability, cost and moderation quality. Each is dated, never edited afterwards, and tied to a version so it can be re-checked later.
One audited my own test coverage, with a deliberately hard question: where is the dangerous untested part? The answer was not flattering. The backend was well tested, but the two front ends were effectively untested. It credited the backend for the right reason too: those tests check real behaviour where trust matters, such as who owns what and who may read whose data, rather than checking status codes and calling that coverage.
| Surface | At audit (June) | Now (August) |
|---|---|---|
| Backend | 85 files | 171 files / 4,862 tests |
| Candidate frontend | ~5 files | 63 files |
| Admin console | 0 behavioural | 45 E2E tests |
The E2E suite starts its own separate backend with a throwaway database, so it never touches production data.
Finding and fixing are separate roles. Test work runs on its own, with written limits: it may read anything, run anything and write tests, but may not change product code. Its standing rule: if you find a bug while writing tests, describe it clearly and stop. That removes the strongest temptation in QA, quietly patching what you found, so that neither the bug nor the reasoning is ever written down.
Important reviews get a second, hostile pass. I re-run them with one instruction: assume the first pass was careless and prove it wrong. The most recent produced 21 corrections. That pass also caught a bug in its own tooling: two stages used different path formats, so seven corrections silently failed to apply, noticed only because the totals did not add up. The rule I wrote afterwards: a step that merges results needs a real check, not a hope, and when two verdicts disagree, keep the stricter one.
What you get
A real entry from my log, with details removed. Note the last section: every report says what the fix does not cover. Writing the limits down stops a one-line fix quietly turning into a rewrite.
Whether a dead listing is caught depends on how many characters are in the employer's name. Detection is a coin flip decided by something unrelated.
The listing is marked as expired. A success response, a redirect to a different page, an error flag in the address, and a title unrelated to the job all mean the same thing: gone.
The listing is kept, and the dead link stays visible to users. With a shorter employer name and nothing else changed, the same listing is correctly expired.
Production backend. Reproduced separately from a local machine against the same addresses.
The check for "is this part a job identifier?" says yes if the part contains a digit or is 12 or more characters, a rule meant for sites that use
words as identifiers. But these addresses look like /{employer}/jobs/{id}, and employer names are
often longer than 12 characters. So the rule decides the job has moved, when what survived
was the company name, which is exactly where a removed job redirects. Separately, the
error flag in the address was never read, because everything after the path is thrown
away first.
Dead links stay visible to users, silently. Measured: roughly one in five affected listings has an employer name long enough to hide the problem, mostly in one source.
Require the most specific identifier to survive: prefer parts with digits when there are any, and use the length rule only when there are none. Also treat an error flag on the final page as the site saying "removed", on an exact name match, and only when the path has changed too.
Because this turns a "keep" into an "expire", it was replayed against the real data before shipping. Result: 3 newly expired, each confirmed dead by hand, 0 wrongly changed, 0 side effects.
It does not change how sources that block checking altogether are handled. That is a separate bug, and combining the two would produce a fix nobody can review. It does not relax the cautious default elsewhere either: expiring too much silently deletes real listings, the more expensive mistake.
Toolbox
How I work
No long warm-up. You get findings in the first week.
I walk through your key flows and rank them by money and regulatory risk. You get a one-page map of where money can leak, before I write a single test case.
Hands-on testing of the highest-risk flows, with repeatable reports in your tracker, in your format. Real bugs in week one, not a test plan you have to read.
Organised test sets for the flows that must never break, plus automation where it pays for itself. Every regression test is proven to fail on the broken code before I trust it.
Updates your managers can read as they are: what shipped, what is at risk, and what I could not check and why. I never report a green I cannot defend.
Contact
iGaming platforms, payment and KYC flows, and quality in systems that move real money. If you are hiring, building a QA team, or want a second opinion on a payment or compliance flow, email is the fastest way to reach me.