l0grisk intelligence · english

// analysis

When AI asks you to pay: examining the performance claims

Illustration for the analysis: When AI asks you to pay: examining the performance claims

An investigation into AI debt-collection claims: receipts, costs, satisfaction, comparison groups and what happens after a repayment agreement.

dated revision: September 24, 2026French originalprimary sourcesno tracker

When AI asks you to pay · Part 2

Read Part 1: who decides to send the reminder?.

On 7 April 2025, Intrum reported that its Ophelos platform had increased amicable recovery rates in the Netherlands by 25%, while reducing the cost to collect by 22%. These were described as early results. The announcement did not provide a protocol that would allow a reader to reproduce the comparison. [1]

The promise deserves a fair hearing. Collecting a valid debt sooner, with less effort and at a lower cost, can benefit both the creditor and the person who owes the money. Yet apparently similar claims can describe very different improvements: more replies, an initial payment, an agreed instalment plan or a debt actually settled.

A system might be excellent at securing commitments without tracking whether they are kept. It could also provide a valuable service by identifying an unjustified demand, leaving the creditor with less money collected. To assess automated debt collection, the inquiry has to continue after the conversation ends.

This documentary investigation draws on company publications, original research and supervisory material available on 24 September 2026. Findings from other countries are not presented as measurements of the French market. No individual case file or production-system test is used.

What sits behind a recovery claim

A brochure distributed by Intrum Belgium under the title AI in Credit Management 2025 reports a 6% improvement in collection results over 120 days, a 13% reduction in collection costs over that period and a satisfaction score rising from 4.21 to 4.38. Its results page gives no calendar dates, group sizes or precise construction of the comparison. Publication through the Belgian business does not, by itself, identify the population being measured. [2]

These numbers do not contradict the Dutch results. Nor do they establish a decline in performance from 25% to 6%. Nothing demonstrates that the same measure, debts or observation period are involved.

Before explaining a difference, it is necessary to establish what is being counted. A rate based on the number of cases producing any payment is different from a rate based on the share of the original amount owed that has been recovered. A few partial payments could improve the first without materially changing the second. Both measures can be useful; they answer different questions.

Time matters too. Over a short horizon, a system might principally bring forward payments that would otherwise have arrived later. That cash-flow benefit has value to a creditor. It does not establish that the final amount recovered will be higher. Distinguishing acceleration from additional recovery requires a sufficiently long follow-up.

Even a closed case needs a definition. Full repayment, a commercial write-off, an upheld dispute, a transfer to another provider and the end of a collection mandate should not disappear into a single category of “resolution”. Doing so would discard the explanation at the very moment the result is counted.

Three different measures An engagement measures a response; a receipt measures money received; verified resolution concerns the outcome of an agreement or dispute. These are not mandatory stages: an upheld dispute may close a case without payment. 01 / READING THE RESULTS What is being counted CONTACT A response secured An open, a reply or a recorded commitment. PAYMENT Money received Receipts reconciled with reversals and refunds. CASE OUTCOME A verified resolution Agreement kept or dispute handled; outcome applied. An unfounded demand may be withdrawn without payment.
l0g analytical framework, without numerical data or a mandatory sequence. An improvement must be tied to the outcome actually measured. Methodological reference: FCA review, 27 July 2026 [12].

Keeping the starting population in view

A robust comparison begins with a cohort: cases entering the process during a defined period and then followed for the same length of time. This prevents newly assigned debts from being compared with debts that have already received months of attention.

The starting group matters as much as the final result. If automation only handles people who can be reached, recent arrears or straightforward amounts, its performance does not describe every case assigned to the collector. Results within that scope might be excellent. The scope still needs to remain visible.

It would be particularly misleading to remove cases handed to a person after a difficulty, then compare the remaining automated cases with the human team’s results. One side would retain the cases it could handle; the other would inherit exceptions as well. A comparison of end-to-end approaches should follow the original groups even when responsibility changes.

That requires a record of entries, exits and the reasons for them. Payments must then be reconciled, reversals and refunds deducted, and principal distinguished from fees. The treatment of interest, concessions and disputed claims needs to be explained. Without that accounting, a change in money recovered could describe something other than a better service.

Results should also be reported as levels as well as changes, with an indication of uncertainty. A relative increase alone says nothing about the starting rate or the number of people affected. A striking percentage might involve few cases; a modest difference might affect a large population.

None of this requires publishing identities or private conversations. Properly defined aggregate statistics, subject to independent review, would already make the claims substantially easier to verify.

Separating the model from the modernisation

Intrum’s own 2025 annual report makes a relevant distinction. On page 19, its chief financial officer separates the savings available from standardising processes from those AI might deliver gradually. This is a discussion of the group’s transformation, not an evaluation of a particular experiment. [3]

Consider a hypothetical implementation: a new portal makes payment easier, staff change their routing rules and a model chooses when reminders are sent. Higher collections could make the project a success. Attributing the entire improvement to the model would still be premature.

The right comparison depends on the question. A creditor considering a complete service may want to know how it performs against the previous arrangement. Establishing AI’s contribution requires asking what would have happened with the same modernised journey, but without the particular AI function being assessed. The project’s cost and the quality of the alternative belong in that calculation.

Where feasible and properly governed, random assignment of comparable cases helps prevent operators from selecting treatments according to case difficulty. It does not remove the need to disclose what differs between groups, retain transferred cases in the analysis or follow outcomes beyond the first payment.

Cost claims demand similar care. Cost per contact can fall while the number of contacts rises. Cost per euro recovered can decline because receipts increase, without an equivalent reduction in total spending. Licences, computing, oversight, integration and subsequent human work need to be included within a stated accounting boundary.

Treating every saving as suspect would be equally unhelpful. Less administrative work and an easier process can serve both parties. The shared benefit needs to be observed, rather than inferred solely from the operator’s lower expenses.

Field experiments follow the money beyond the handover

A paper by James Choi and his co-authors, issued by the NBER in April 2025, studies a Chinese lender using data from April 2021 to December 2023. The analysis includes cases entering collection before December 2022, allowing at least a year of follow-up. In its randomised experiment, some cases are assigned to AI on days two to five past due, then to people from day six; others receive human calls from day two. Cases initially assigned to AI yield less under the authors’ measure. The gap narrows sharply after the human handover but remains at 360 days. [6] [7]

The measure also accounts for payment timing. The authors discount receipts, assigning less value to money received later, then divide the total by the initial overdue amount. That measure must not be retold as a simple rate of cases paid off. The protocol and tables are available in the author version dated 30 March 2025. [7]

This is not a test of generative agents deployed in France in 2026. Its relevance is the need to assess the whole journey: obtaining a promise or making a successful human handover is not the final outcome. Yale’s account places the deployment before the widespread release of today’s generative tools. [8]

Other findings prevent this weaker collection result from being generalised. A study by Fang and co-authors published online by Manufacturing & Service Operations Management on 7 May 2026 reports two randomised field experiments. Its abstract says undisclosed AI agents outperform people under some context-appropriate emotional instructions, but perform worse under inappropriate instructions. Only the abstract and publication record could be examined here, so the claimed percentage gains are not reproduced. [9]

These studies assess different systems and workflows. They do not settle a contest between “AI” and “humans” in general. They make it necessary to specify the instructions, population, follow-up period and interaction conditions before applying a result to another service.

Disclosure is one such condition. In the European Union, Article 50 of the AI Act has applied since 2 August 2026. For direct interactions within its scope, providers must design systems to inform people that they are interacting with AI, subject notably to the exception where this is obvious. An experiment involving undisclosed AI is therefore not a ready-made deployment recipe. [10]

Examining PAIR’s claim about emotional stress

On 11 February 2026, PAIR Finance presented a company-commissioned study as showing a 36% reduction in emotional stress associated with debt collection. Its calculation uses reported proportions of participants feeling judged: 61% with a human and 39% with AI. The data were collected between August and October 2025. [4]

The arithmetic is straightforward: the stated decline is 22 percentage points, or approximately 36% of the initial level. That checks the conversion between figures, not how the underlying proportions were constructed. Feeling judged is also one particular aspect of an experience. It should not become a comprehensive measure of stress or evidence of improved financial health.

The preprint by Goetze, Clajus and Stricker involves 3,514 people randomly assigned to two groups. Each participant read one conversation scenario. They did not speak to an agent or repay a debt as part of the test. [5]

The appendix reveals an important design feature. The human scenario specifies a ten-minute wait, the AI scenario a ten-second wait. The human also comments that earlier contact would have been preferable; the AI thanks the person for being candid. Several features change together. [5]

The experiment compares two complete scenarios Participants read one of two scenarios. The human scenario specifies ten minutes of waiting and a remark about earlier contact; the AI scenario specifies ten seconds and thanks the person for being candid. The comparison changes several features at once, not only the interlocutor’s identity. 02 / READING THE DESIGN Two scenes to read One scenario per participant HUMAN SCENARIO 10 minutes of stated waiting time A remark about not getting in touch earlier. AI SCENARIO 10 seconds of stated waiting time Thanks for being candid. Waiting and wording also change. The effect of AI identity alone is not isolated.
Fictional research scenarios, not measurements of a production service. Data collected August-October 2025. Summary of differences in Table A1, pp. 14-15, Goetze and co-authors; methodological interpretation by l0g. [5]

Random assignment makes it possible to compare reactions to those two bundles of features. It does not isolate the effect of the interlocutor’s artificial identity. Would a prompt, non-judgmental human response have produced the same reaction? The design does not supply that comparison.

That does not make the study useless. A digital service that is genuinely more available and less guilt-inducing could be an improvement. The remaining questions are whether the deployed service reproduces those characteristics and whether the effect holds among people dealing with an actual debt. Assessing a written scene does not establish whether an instalment plan is affordable or whether a person can obtain redress when an interaction goes wrong.

Readers should know that PAIR commissioned the research and that two co-authors are affiliated with the company. Those links do not invalidate the observations. Publishing the protocol is precisely what makes scrutiny possible. The study nevertheless does not independently validate the supplier’s full set of commercial performance claims. [4] [5]

Who never appears in the satisfaction score

Positive feedback is useful evidence. It also needs to be placed alongside the people whose views were not recorded.

On page 75 of its annual report, Intrum describes a post-interaction SMS survey. The group explicitly acknowledges selection and non-response limitations: the survey reaches people who have had an interaction, can be contacted through that channel and choose to reply. Respondents may therefore not represent the whole contacted population. [3]

A higher average could reflect better service, but also a change in who responds. Distinguishing the two requires invitation counts, response rates and a description of the population covered. An improvement among portal users does not reveal the experience of people who never entered the portal.

Call volumes need the same caution. Intrum’s Belgian brochure reports a 17% reduction in incoming traffic, attributed to self-service adoption. [2] The explanation is plausible: someone who can resolve a matter alone has less reason to call. Verification should nevertheless distinguish successful self-service from abandoning an attempt to get help.

An avoided contact, an abandoned call and a resolved request are not interchangeable outcomes. Looking at them together would help establish whether lower traffic means less need for assistance or more difficulty obtaining it. These are competing possibilities to investigate, not established incidents at the company.

The US Consumer Financial Protection Bureau documented complaints about repetitive automated journeys and difficulties reaching human support in a June 2023 report on financial-service chatbots. It supplies situations worth monitoring, not a failure rate for today’s European platforms. [11]

What remains after the payment

On 27 July 2026, the UK’s Financial Conduct Authority published a review of consumer-outcome monitoring. It highlights the limits of using activity measures as substitutes for actual effects and the value of combining different sources of evidence. The review concerns firms subject to the UK Consumer Duty. It is neither a review of French debt collection nor a statement of French law. [12]

Applied to this investigation, that distinction means examining receipts, the fate of agreements and corrections together. An accepted instalment plan should be followed to its end, separating plans still running from those that have reached their scheduled completion. Looking only at finished plans would discard part of the story.

A dispute requires a different test: was the challenge reviewed, and was the outcome implemented? Where an error is accepted, verification should extend to the corrected amount, any refund and the cessation of action that no longer has a basis. A courteous response does not establish those things. Neither does an account balance on its own.

For someone in financial difficulty, the ability to keep an agreement needs to be considered alongside other commitments. Paying one debt could leave rent or another essential bill unpaid. That risk supports proportionate assessment; it does not justify unlimited collection of intimate information or an assumption that every payment causes hardship.

A strong validation package would bring together cohorts defined before analysis, outcomes measured over several horizons and indicators covering errors and redress. Scope choices, version changes and exclusions would remain visible. Difficulties arising after a handover to a human would stay in the assessment of the initial journey.

It would also examine benefits for people who do not pay: an unfounded demand withdrawn, a referral to suitable support or revised arrangements when circumstances change. Collections might remain unchanged while the handling of the case improves substantially.

What the evidence supports

The publications examined describe commercial gains and credible opportunities to simplify collection. They do not establish a market-wide causal effect of AI in French debt collection, or an average benefit for the people concerned. Local comparisons, psychological scenarios and overseas field experiments do not fill that gap simply by being added together.

The strongest conclusion concerns how to verify the promises. The original population, handover decisions and events after the initial payment must remain in view. Money collected needs to be assessed alongside full costs and the handling of difficulties. Satisfaction has to be distinguished from silence.

A debt-collection AI should be evaluated through to the outcome of the agreement or dispute. Without that follow-up, we may know that it secured a response. We will not yet know whether it helped resolve the situation.


Method and limitations

Company documents establish what their authors claim, not an independent replication of the results. Relevant pages and appendices are identified below. The author version of Choi and co-authors’ paper was consulted. Discussion of Fang and co-authors’ study is limited to the publisher’s abstract and metadata. Statistical analyses were not replicated: individual-level data and replication code were not used in this investigation. No interviews, requests for comment sent to the parties or live-system tests were conducted. Questions requiring an answer remain documentary limitations, not established violations.

Sources and scope

[1] Intrum · Proven impact: Our AI platform delivers results in both implementation and innovation. 7 April 2025. Early Dutch results, not a French estimate. Cohorts and comparison protocol are not detailed.

[2] Intrum Belgium · AI in Credit Management 2025. Brochure identified as 2025; exact release date unconfirmed. Page 5: collection and cost results over 120 days, satisfaction and incoming traffic. Population not detailed.

[3] Intrum · Annual Report 2025. Financial year 2025. Page 19: process savings and AI’s gradual contribution. Page 75: SMS survey design and limitations. Group and local populations are not interchangeable.

[4] PAIR Finance · Étude européenne : l’IA réduit le stress émotionnel lié au recouvrement de créances de 36 %. 11 February 2026; data collected August-October 2025. Source of the 36% claim and reported 61% / 39% proportions. The release describes the commissioning and co-author affiliations.

[5] Minou Goetze, Sebastian Clajus, Stephan Stricker · AI in Debt Collection: Estimating the Psychological Impact on Consumers. Preprint v1, listed as submitted on 19 January 2026. Methods p. 4; results p. 6; Table A1 scenarios pp. 14-15. Participants read scenarios rather than experiencing live conversations.

[6] James J. Choi, Dong Huang, Zhishu Yang, Qi Zhang / NBER · How Good is AI at Twisting Arms? Experiments in Debt Collection: Working Paper 33669. April 2025. NBER record and abstract; the full paper was read through the author version below. This is a working paper.

[7] James J. Choi and co-authors, version linked from his Yale page · How Good is AI at Twisting Arms? Experiments in Debt Collection. Author version dated 30 March 2025, linked from the author’s Yale research page. Printed pagination: data and design pp. 10-17; Table 4 p. 48. Outcomes are discounted receipts scaled by the initial overdue balance.

[8] Yale School of Management / Yale Insights · Can AI Replace Human Debt Collectors?. 20 May 2025. University account of its researcher’s work, used to situate the generation of systems studied. Units and empirical details were checked in the original paper.

[9] Zheng Fang, Yuqian Chang, Xueming Luo, Qingsheng Wu, Jaakko Aspara / INFORMS · Artificial Intelligence, Emotional Labor, and Service Operations. Online publication 7 May 2026; accepted 1 April 2026. Abstract and metadata only. Numerical gains are omitted because the full text and comparison bases could not be examined.

[10] European Commission · Transparency obligations under Article 50 of the AI Act. FAQ updated 24 July 2026. Article 50 applies from 2 August 2026; explains the scope and exceptions for disclosure of AI interactions.

[11] Consumer Financial Protection Bureau · Chatbots in consumer finance. 6 June 2023. US analysis of financial chatbots and complaints, including response loops and access to support. It does not supply a French failure rate.

[12] Financial Conduct Authority · Outcomes monitoring: good practice and areas for improvement. 27 July 2026. Sections 1, 5 and 6: activity versus outcomes, journey monitoring and combining indicators. A UK reference, not French law.

Documents consulted on 24 September 2026. Publication dates and observation periods are distinguished where known.

This analysis is not investment advice.

// cite this analysis

l0g, “When AI asks you to pay: examining the performance claims”, l0g.fr, published September 24, 2026, updated September 24, 2026, https://l0g.fr/en/analysis/ai-debt-collection-2-performance-claims/


$ cd ../analysis