Machine Learning · Architecture
An ML Score With Nothing to Train On:
Plug It In, Don't Let It Decide
A machine-learning risk score can sit next to a rule-based refund decision even when there are no real outcomes to train it on. Put it behind an interface, pick the implementation by config, fail open when it breaks, and leave the decision to the domain rule. The scorer below was wired that way, and has since been removed.
RefundFraudRiskScorer described below has since been removed too. It didn't share the flaw covered in The Fraud Signal That Trusted the Fraudster (the requester's own history isn't something they can rewrite on request); it went as a separate simplification decision made at the same time. RefundEligibilityService now carries no fraud-risk judgment of any kind. The source link at the bottom of this post now points to the last commit where the file still existed.
The problem was adding a fraud-risk score to a refund decision with no history of real fraud reviews to learn from. In my example project that implements the same backend design in five languages side by side, refund approval was decided by a rule-based Domain Service, RefundEligibilityService. The scorer I put next to it, RefundFraudRiskScorer, was a hand-rolled logistic regression reading structured numbers from the requester's history: refund count, rejection rate, amount ratio, and time since payment. By default there was no LLM and no external API, just four features and a sigmoid.
It was the second of two Technical Services feeding that Domain Service, and the two were deliberately different shapes of "machine learning." The other, RefundReasonClassifier, was an LLM reading what the customer wrote as a reason. Neither one got the final vote.
Put the Score Behind an Interface
interface RefundFraudRiskScorer {
fun score(features: RefundRiskFeatures): Double
}Two classes implemented it, selected by config rather than by the caller. RequestRefundService depended only on the interface and never knew which one was live.
Features From the Requester's Own History
Everything the model saw came from the requester's own history, assembled by the Application layer from the Payment and Refund Aggregates plus a repository summary query:
val mlFraudRiskScore =
refundFraudRiskScorer.score(
RefundRiskFeatures(
refundCountLast30Days = refundSummary.count.toInt(),
rejectedRefundCountLast30Days = rejectedRefundSummary.count.toInt(),
refundToPaymentAmountRatio = refund.amount.toDouble() / payment.amount.toDouble(),
minutesSincePayment =
Duration.between(payment.createdAt, LocalDateTime.now())
.toMinutes()
.coerceAtLeast(0)
.toDouble(),
),
)With No Data, Train on a Placeholder
An example project has no real users, so there was no real historical fraud-review outcome to train against. The native implementation trained itself once, at construction, against a synthetic seeded dataset and a deliberately simple ground-truth rule:
private fun generateTrainingData(): List<TrainingExample> {
val random = Random(TRAINING_SEED)
return (0 until TRAINING_EXAMPLE_COUNT).map {
val refundCountLast30Days = random.nextInt(8)
val rejectedRefundCountLast30Days = random.nextInt(4)
val refundToPaymentAmountRatio = random.nextDouble()
val minutesSincePayment = random.nextDouble() * 43200
val riskScore =
refundCountLast30Days * 0.15 +
rejectedRefundCountLast30Days * 0.3 +
refundToPaymentAmountRatio * 0.4 +
maxOf(0.0, 1 - minutesSincePayment / 1440) * 0.3
val label = if (riskScore > 1.1) 1.0 else 0.0
TrainingExample(/* ... */ label = label)
}
}Plain batch gradient descent, four weights plus a bias, no ML library:
private fun trainLogisticRegression(examples: List<TrainingExample>): LogisticModel {
val weights = DoubleArray(FEATURE_COUNT)
var bias = 0.0
repeat(EPOCHS) {
val weightGradients = DoubleArray(FEATURE_COUNT)
var biasGradient = 0.0
for (example in examples) {
val vector = toVector(example.features)
var z = bias
for (i in vector.indices) z += vector[i] * weights[i]
val error = sigmoid(z) - example.label
for (i in vector.indices) weightGradients[i] += error * vector[i]
biasGradient += error
}
for (i in weights.indices) weights[i] -= (LEARNING_RATE * weightGradients[i]) / examples.size
bias -= (LEARNING_RATE * biasGradient) / examples.size
}
return LogisticModel(weights, bias)
}The fixed random seed mattered here: the generated dataset, and therefore the trained weights, came out identical on every run. The model was explicitly a stand-in; the interface was what mattered, not its predictive power.
Swap It by Config, Fail Open When It Breaks
The same native/HTTP toggle already used for the LLM classifier showed up here too. A config property picked between an in-process computation and a call to the shared services/fraud-risk-scorer microservice:
@ConfigurationProperties(prefix = "fraud-scorer")
data class FraudScorerProperties(
val mode: String = "native",
val baseUrl: String = "http://localhost:8000",
) {
val isHttpMode: Boolean get() = mode == "http"
}The HTTP implementation failed open. Any network error, non-2xx, or malformed response returned a score of 0.0 rather than blocking the refund:
override fun score(features: RefundRiskFeatures): Double =
try {
val response = httpClient.send(buildRequest(features), HttpResponse.BodyHandlers.ofString())
if (response.statusCode() !in 200..299) FALLBACK_SCORE else parseScore(response.body()) ?: FALLBACK_SCORE
} catch (e: Exception) {
// A scoring failure is a technical-infrastructure concern, not a domain error — it must
// never block a refund request. Swallow it here at the boundary and fall back.
FALLBACK_SCORE
}Let the Domain Rule Decide
RefundEligibilityService took both signals as independent values, each with its own threshold, and neither Technical Service knew the other existed:
companion object {
private const val FRAUD_RISK_REJECTION_THRESHOLD = 0.7 // from RefundReasonClassifier (LLM)
private const val ML_FRAUD_RISK_REJECTION_THRESHOLD = 0.8 // from RefundFraudRiskScorer (history model)
}
fun evaluate(payment: Payment, refund: Refund, classification: RefundReasonClassification, mlFraudRiskScore: Double): RefundDecision {
// ...
if (mlFraudRiskScore >= ML_FRAUD_RISK_REJECTION_THRESHOLD) {
return RefundDecision(approved = false, reason = "This refund pattern was flagged as high risk by the fraud-risk model and requires manual review.")
}
return RefundDecision(approved = true)
}The Domain Service was the only place both numbers met, and the only place that decided what they meant.
A History-Based Score Makes Tests Share State
Adding a history-aware scorer to an E2E suite with shared test fixtures created a deterministic failure in unrelated tests, not a flaky one. Multiple test methods reusing the same owner ID against a Testcontainers Postgres instance (no per-test reset) meant later tests inherited rejected-refund history from earlier ones, pushing the native score past the 0.8 threshold and misclassifying a legitimately valid refund as high-risk.
The two language implementations that hit this fixed it two different ways, and it's worth naming both rather than claiming one shared technique. The java-springboot implementation forced its entire E2E suite into HTTP mode against an unreachable address, so scoring deterministically fell back to 0 for every test. The nestjs implementation instead left native scoring live for the rest of the suite and gave only the one affected test its own dedicated owner ID. That's a narrower fix for the same underlying cause.
It's worth naming the difference: a flaky test fails unpredictably for reasons unrelated to the code under test. This failure happened every time, in the same order, for the same reason: accumulated state from earlier tests changing the input to a later one. That's a test-isolation bug wearing a "flaky test" costume, and it's worth looking twice before reaching for a retry-on-failure fix instead of an isolation fix.
docs/architecture/domain-service.md (the Technical Service pattern in my example project that implements the same backend design in five languages; the example in this post has since been replaced there, see the update note above) · RefundFraudRiskScorerNativeImpl.kt (the training and scoring code as it existed, pinned to the last commit before removal)