Tax fraud detection for under-reporting declarations using an unsupervised machine learning approach

Daniel De Roux, Boris Pérez, Andrés Moreno, Maria Del Pilar Villamil, César Figueroa

Producción: Capítulo del libro/informe/acta de congresoContribución a la conferenciarevisión exhaustiva

66 Citas (Scopus)

Resumen

Tax fraud is the intentional act of lying on a tax return form with intent to lower one's tax liability. Under-reporting is one of the most common types of tax fraud, it consists in filling a tax return form with a lesser tax base. As a result of this act, fiscal revenues are reduced, undermining public investment. Detecting tax fraud is one of the main priorities of local tax authorities which are required to develop cost-efficient strategies to tackle this problem. Most of the recent works in tax fraud detection are based on supervised machine learning techniques that make use of labeled or audit-assisted data. Regrettably, auditing tax declarations is a slow and costly process, therefore access to labeled historical information is extremely limited. For this reason, the applicability of supervised machine learning techniques for tax fraud detection is severely hindered. Such limitations motivate the contribution of this work. We present a novel approach for the detection of potential fraudulent tax payers using only unsupervised learning techniques and allowing the future use of supervised learning techniques. We demonstrate the ability of our model to identify under-reporting taxpayers on real tax payment declarations, reducing the number of potential fraudulent tax payers to audit. The obtained results demonstrate that our model doesn't miss on marking declarations as suspicious and labels previously undetected tax declarations as suspicious, increasing the operational efficiency in the tax supervision process without needing historic labeled data.

Idioma originalInglés
Título de la publicación alojadaKDD 2018 - Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
EditorialAssociation for Computing Machinery
Páginas215-222
Número de páginas8
ISBN (versión impresa)9781450355520
DOI
EstadoPublicada - 19 jul. 2018
Publicado de forma externa
Evento24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD 2018 - London, Reino Unido
Duración: 19 ago. 201823 ago. 2018

Serie de la publicación

NombreProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining

Conferencia

Conferencia24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD 2018
País/TerritorioReino Unido
CiudadLondon
Período19/08/1823/08/18

Huella

Profundice en los temas de investigación de 'Tax fraud detection for under-reporting declarations using an unsupervised machine learning approach'. En conjunto forman una huella única.

Citar esto