Building Adversarially Robust Predictive Systems for Tabular Data

🧑‍💻 Zhipeng “Zippo” He 何志鹏 @ School of Information Systems
Queensland University of Technology

  • XAMI LAB
  • zhipeng.he@hdr.qut.edu.au
  • github.com/ZhipengHe
  • zhipenghe.me


PhD Final Seminar
October 08, 2025

Supervisory Team

  • A/Prof. Chun Ouyang (QUT)
  • Prof. Alistair Barros (QUT)
  • A/Prof. Catarina Moreira (UTS)

🧭 Outline

💡 Background & Motivation

🎯 Research Problem

🧩 Objective 1 — Foundation of Imperceptibility
Defining imperceptible properties on tabular data.

📊 Objective 2 — Benchmark of Tabular Attacks
Systematic evaluation framework for existing adversarial attacks.

⚙️ Objective 3 — Generative VAE-based Attack
Unified treatment of categorical & numerical features in latent space.

🌟 Conclusion & Future Directions
From conceptual clarity → systematic evaluation → practical robustness.

💡 Background & Motivation

Trustworthy AI: Lawful, Ethical, and Robust

The relation between different aspects of Artificial Intelligence (AI) trustworthiness (B. Li et al. 2023).
Li, Bo, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou. 2023. “Trustworthy AI: From Principles to Practices.” ACM Computing Surveys 55 (9): 1–46. https://doi.org/10.1145/3555803.

What is Robustness in CS & ML ?

In Computer Science (CS), robustness has been more formally defined by the IEEE (IEEE 2008) as:

the degree to which a system or component can function correctly in the presence of invalid inputs or stressful conditions.


In Machine Learning (ML), robustness refers to:

the ability of a model to perform well on unseen or new data or variations in the input.

IEEE. 2008. IEEE Standard Computer Dictionary: A Compilation of IEEE Standard Computer Glossaries.” Piscataway, NJ, USA: IEEE. https://doi.org/10.1109/ieeestd.1991.106963.

What are the Input Variations?

Different sources of input variation that challenge robustness in ML

Adversarial robustness = robustness against worst-case, deliberately crafted perturbations (Szegedy et al. 2014)

Szegedy, Christian, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2014. “Intriguing Properties of Neural Networks.” In 2nd International Conference on Learning Representations (ICLR 2014), edited by Yoshua Bengio and Yann LeCun. https://arxiv.org/abs/1312.6199.

What are Adversarial Attacks?

An adversarial attack is a method to generate adversarial examples.

Adversarial examples are specialised inputs created with the purpose of confusing a neural network, resulting in the misclassification of a given input. These notorious inputs are indistinguishable to the human eye but cause the network to fail to identify the contents of the image. (Goodfellow, Shlens, and Szegedy 2015)

Goodfellow, Ian J, Jonathon Shlens, and Christian Szegedy. 2015. “Explaining and Harnessing Adversarial Examples.” In 3rd International Conference on Learning Representations (ICLR 2015), edited by Yoshua Bengio and Yann LeCun. https://arxiv.org/abs/1412.6572.

What are Adversarial Attacks?

Illustration of the process of adversarial attack.

What are Adversarial Attacks?

For example, when applying Fast Gradient Sign Method (FGSM), we can generate an adversarial example:

Adversarial example generated with Fast Gradient Sign Method (FGSM) (Goodfellow, Shlens, and Szegedy 2015)

Even small, imperceptible perturbations can cause a model to make confidently incorrect predictions.

Goodfellow, Ian J, Jonathon Shlens, and Christian Szegedy. 2015. “Explaining and Harnessing Adversarial Examples.” In 3rd International Conference on Learning Representations (ICLR 2015), edited by Yoshua Bengio and Yann LeCun. https://arxiv.org/abs/1412.6572.

From Computer Vision to Tabular Data

  • Adversarial attacks have been extensively studied in Computer Vision
  • But what about tabular data?

EHR Records

Credit Card Transactions

Network Traffic Logs

Why Tabular Data Matters

Tabular data powers critical real-world systems: finance (Lunghi et al. 2023), healthcare (Sun et al. 2018), cybersecurity (D. Li et al. 2023)

Image (unstructured)
• High-dimensional, homogeneous
• Continuous pixels

Tabular (structured) (Borisov et al. 2024)
• Lower-dimensional, heterogeneous
• Numerical + categorical
• Feature dependencies


Would existing adversarial attacks still work on tabular data without adaptations?

Borisov, Vadim, Tobias Leemann, Kathrin Sebler, Johannes Haug, Martin Pawelczyk, and Gjergji Kasneci. 2024. “Deep Neural Networks and Tabular Data: A Survey.” IEEE Transactions on Neural Networks and Learning Systems 35 (6): 7499–519. https://doi.org/10.1109/TNNLS.2022.3229161.
Li, Deqiang, Qianmu Li, Yanfang (fanny) Ye, and Shouhuai Xu. 2023. “Arms Race in Adversarial Malware Detection: A Survey.” ACM Computing Surveys 55 (1): 1–35. https://doi.org/10.1145/3484491.
Lunghi, Daniele, Alkis Simitsis, Olivier Caelen, and Gianluca Bontempi. 2023. “Adversarial Learning in Real-World Fraud Detection: Challenges and Perspectives.” In Proceedings of the Second ACM Data Economy Workshop (DEC 2023), 27–33. New York, NY, USA: ACM. https://doi.org/10.1145/3600046.3600051.
Sun, Mengying, Fengyi Tang, Jinfeng Yi, Fei Wang, and Jiayu Zhou. 2018. “Identify Susceptible Locations in Medical Records via Adversarial Attacks on Deep Predictive Models.” In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD 2018), edited by Yike Guo and Faisal Farooq. New York, NY, USA: ACM. https://doi.org/10.1145/3219819.3219909.

Adversarial Threat Models

Adversarial attacks can be described along four dimensions (Papernot et al. 2016; Biggio and Roli 2018; Assion et al. 2019; Pelekis et al. 2025):

A taxonomy of adversarial attack threat models. Four dimensions: Influence, Knowledge, Constraints, Goals.
Assion, Felix, Peter Schlicht, Florens Gressner, Wiebke Gunther, Fabian Huger, Nico Schmidt, and Umair Rasheed. 2019. “The Attack Generator: A Systematic Approach Towards Constructing Adversarial Attacks.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW 2019), 1370–79. IEEE. https://doi.org/10.1109/cvprw.2019.00177.
Biggio, Battista, and Fabio Roli. 2018. “Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning.” Pattern Recognition 84: 317–31. https://doi.org/10.1016/j.patcog.2018.07.023.
Papernot, Nicolas, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami. 2016. “The Limitations of Deep Learning in Adversarial Settings.” In 2016 IEEE European Symposium on Security and Privacy (EuroS&p 2016), 372–87. IEEE. https://doi.org/10.1109/eurosp.2016.36.
Pelekis, Sotiris, Thanos Koutroubas, Afroditi Blika, Anastasis Berdelis, Evangelos Karakolis, Christos Ntanos, Evangelos Spiliotis, and Dimitris Askounis. 2025. “Adversarial Machine Learning: A Review of Methods, Tools, and Critical Industry Sectors.” Artificial Intelligence Review 58 (8): 1–87. https://doi.org/10.1007/s10462-025-11147-4.

When to Deploy Adversarial Attacks?

Privacy Attack (training-time) out of scope

Poison Attack (training-time) out of scope

Evasion Attack (test-time)


Why focus on test-time attacks? (Pelekis et al. 2025)

  • Privacy attacks: Extract sensitive training data to break the model
  • Poison attacks: Corrupt model behavior during training to break the model
  • Evasion attacks: Test adversarial robustness of deployed models against adversarial examples
Pelekis, Sotiris, Thanos Koutroubas, Afroditi Blika, Anastasis Berdelis, Evangelos Karakolis, Christos Ntanos, Evangelos Spiliotis, and Dimitris Askounis. 2025. “Adversarial Machine Learning: A Review of Methods, Tools, and Critical Industry Sectors.” Artificial Intelligence Review 58 (8): 1–87. https://doi.org/10.1007/s10462-025-11147-4.

What Access do Adversarial Attacks have?

White-Box Attacks ✅

Grey-Box Attacks ❌

Black-Box Attacks ❌

  • Full access to model and dataset knowledge
  • Partial access to model and dataset knowledge
  • No access to model knowledge; but can query the model

White-box attacks represent the worst-case scenarios (Papernot et al. 2016)

👉 If a model is robust under white-box attacks, it can handle any other type of attacks.

Papernot, Nicolas, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami. 2016. “The Limitations of Deep Learning in Adversarial Settings.” In 2016 IEEE European Symposium on Security and Privacy (EuroS&p 2016), 372–87. IEEE. https://doi.org/10.1109/eurosp.2016.36.

How to Constrain the Adversarial Attacks?

Bounded attacks

  • Perturbations restricted by a budget (\(\lVert \delta \rVert \leq \epsilon\))
  • Keep perturbed inputs “close” to original

Bounded perturbations (\(\ell_\infty\) ball)

Unbounded attacks

  • Minimal perturbation to cross decision boundary
  • Often harder to compute but insightful

Unbounded perturbations (decision boundary)

What are the Goals of Adversarial Attacks?

Confidence Reduction ❌

Nontargeted Misclassification ✅

Targeted Misclassification ❌

Source/Target Misclassification ❌

Given a credit score classification systems \(f: \mathbb{X}\to \mathbb{Y}\) with three classes \(\mathbb{Y} = \{\)Poor, Standard, Good\(\}\), select an input instance \(\boldsymbol{x}\in\mathbb{X}\) to craft a adversarial example \(\boldsymbol{x}^{adv}\):

  1. Confidence Reduction: Attack not success: \(\boldsymbol{x}\) (Poor, 90% confidence) → \(\boldsymbol{x}^{adv}\) (Poor, 50% confidence)
  2. Nontargeted Misclassification: Attack success:\(\boldsymbol{x}\) (any class, e.g., Poor) → \(\boldsymbol{x}^{adv}\) (a different class, Standard or Good)
  3. Targeted Misclassification: Attack success: \(\boldsymbol{x}\) (any class, e.g., Poor or Standard) → \(\boldsymbol{x}^{adv}\) (specific class Good)
  4. Source/Target Misclassification: Attack success:\(\boldsymbol{x}\) (specific class Poor) → \(\boldsymbol{x}^{adv}\) (specific class Good)

Summary of Research Scope

A taxonomy of adversarial attack threat models. Four dimensions: Influence, Knowledge, Constraints, Goals. Highlighted boxes show research scope.

Benchmarks in Adversarial Attacks

Most existing benchmarks focus on:

The only tabular benchmark focuses on a single attack method (Simonetto, Ghamizi, and Cordy 2024), leaving comprehensive tabular evaluation underexplored.

Cinà, Antonio Emanuele, Jérôme Rony, Maura Pintor, Luca Demetrio, Ambra Demontis, Battista Biggio, Ismail Ben Ayed, and Fabio Roli. 2025. AttackBench: Evaluating Gradient-Based Attacks for Adversarial Examples.” In Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence (AAAI-25), edited by Toby Walsh, Julie Shah, and Zico Kolter, 2600–2608. Association for the Advancement of Artificial Intelligence (AAAI). https://doi.org/10.1609/aaai.v39i3.32263.
Croce, Francesco, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. 2021. RobustBench: A Standardized Adversarial Robustness Benchmark.” In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS Datasets and Benchmarks 2021), edited by Joaquin Vanschoren and Sai-Kit Yeung. Vol. 1. Curran Associates, Inc. https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/a3c65c2974270fd093ee8a9bf8ae7d0b-Abstract-round2.html.
Hingun, Nabeel, Chawin Sitawarin, Jerry Li, and David Wagner. 2023. REAP: A Large-Scale Realistic Adversarial Patch Benchmark.” In 2023 IEEE/CVF International Conference on Computer Vision (ICCV 2023), 4640–51. IEEE. https://doi.org/10.1109/ICCV51070.2023.00428.
Siddiqui, Shoaib Ahmed, Andreas Dengel, and Sheraz Ahmed. 2020. “Benchmarking Adversarial Attacks and Defenses for Time-Series Data.” In 27th International Conference on Neural Information Processing (ICONIP 2020), 12534:544–54. Lecture Notes in Computer Science. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-63836-8_45.
Simonetto, Thibault, Salah Ghamizi, and Maxime Cordy. 2024. “Towards Adaptive Attacks on Constrained Tabular Machine Learning.” In ICML 2024 Workshop on the Next Generation of AI Safety. OpenReview.net. https://openreview.net/forum?id=DnvYdmR9OB.
Zheng, Qinkai, Xu Zou, Yuxiao Dong, Yukuo Cen, Da Yin, Jiarong Xu, Yang Yang, and Jie Tang. 2021. “Graph Robustness Benchmark: Benchmarking the Adversarial Robustness of Graph Machine Learning.” In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS Datasets and Benchmarks 2021), edited by Joaquin Vanschoren and Sai-Kit Yeung. Curran Associates, Inc. https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/6cdd60ea0045eb7a6ec44c54d29ed402-Abstract-round2.html.

Overall Problem

Researchers have been working on making models more robust to adversarial attacks on unstructured data. However, researchers have rarely investigated adversarial robustness for tabular data.

How can one construct predictive models that are robust to adversarial attacks for tabular data?

Adversarial Attacks on Tabular Data

From the definition of adversarial examples, we have two central challenges:

Ballet, Vincent, Xavier Renard, Jonathan Aigrain, Thibault Laugel, Pascal Frossard, and Marcin Detyniecki. 2019. “Imperceptible Adversarial Attacks on Tabular Data.” arXiv [Stat.ML]. http://arxiv.org/abs/1911.03274.
Carlini, Nicholas, and David Wagner. 2017. “Towards Evaluating the Robustness of Neural Networks.” In 2017 IEEE Symposium on Security and Privacy (SP 2017), 39–57. IEEE. https://doi.org/10.1109/sp.2017.49.
Chernikova, Alesia, and Alina Oprea. 2022. FENCE: Feasible Evasion Attacks on Neural Networks in Constrained Environments.” ACM Transactions on Privacy and Security 25 (4): 1–34. https://doi.org/10.1145/3544746.
Goodfellow, Ian J, Jonathon Shlens, and Christian Szegedy. 2015. “Explaining and Harnessing Adversarial Examples.” In 3rd International Conference on Learning Representations (ICLR 2015), edited by Yoshua Bengio and Yann LeCun. https://arxiv.org/abs/1412.6572.
Kurakin, Alexey, Ian Goodfellow, and Samy Bengio. 2016. “Adversarial Examples in the Physical World.” arXiv [Cs.CV]. http://arxiv.org/abs/1607.02533.
Madry, Aleksander, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. “Towards Deep Learning Models Resistant to Adversarial Attacks.” In 6th International Conference on Learning Representations (ICLR 2018). OpenReview.net. https://openreview.net/forum?id=rJzIBfZAb.
Moosavi-Dezfooli, Seyed-Mohsen, Alhussein Fawzi, and Pascal Frossard. 2016. DeepFool: A Simple and Accurate Method to Fool Deep Neural Networks.” In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), 2574–82. IEEE. https://doi.org/10.1109/cvpr.2016.282.

Imperceptibility in Computer Vision

The same example also shows another key property: adversarial perturbations are designed to be imperceptible.

In Computer Vision, adversarial perturbations must be “indistinguishable to the human eye” (Goodfellow, Shlens, and Szegedy 2015) — the adversarial image should look identical to the original.

Goodfellow, Ian J, Jonathon Shlens, and Christian Szegedy. 2015. “Explaining and Harnessing Adversarial Examples.” In 3rd International Conference on Learning Representations (ICLR 2015), edited by Yoshua Bengio and Yann LeCun. https://arxiv.org/abs/1412.6572.

Imperceptibility in Computer Vision

This is formalised using bounded perturbations under norms (e.g., \(\ell_\infty\)) (Madry et al. 2018). These constraints limit the size of pixel-level changes.

Bounded perturbations under norms (e.g., \(\ell_\infty\) norm)
Madry, Aleksander, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. “Towards Deep Learning Models Resistant to Adversarial Attacks.” In 6th International Conference on Learning Representations (ICLR 2018). OpenReview.net. https://openreview.net/forum?id=rJzIBfZAb.

Imperceptibility in Tabular Data: A Challenge

The failure of image-based imperceptibility concepts for tabular data.

👉 CV-style definitions of imperceptibility (pixel norms) do not transfer to tabular data.

Redefining Imperceptibility for Tabular Data

Goal: adversarial examples should not just small perturbations, but also remain realistic.

Imperceptibility in tabular data means perturbations must preserve:

  • Statistical consistency → numerical features follow realistic distributions

  • Semantic validity → categorical values remain logically consistent

  • Structural integrity → dependencies between features are not violated

🎯 Research Problem

Research Question & Objective 1

RQ1: What properties can be used to define the imperceptibility of adversarial attacks on tabular data?

Unlike images, tabular datasets require definitions that account for

  • statistical consistency
  • semantic validity
  • structural integrity

This question establishes the conceptual foundations of imperceptibility.

Objective 1: Identify and formalise the set of properties of imperceptible adversarial attacks on tabular data.

Research Question & Objective 2

RQ2: How do existing adversarial attacks perform under both effectiveness and tabular-specific imperceptibility criteria?

Two dimensions:
1. Ability to misclassify predictive models (effectiveness)
2. Extent to which perturbations achieve minimal perturbation and statistical consistency (imperceptibility)

Sub-questions:

  • RQ2.1: How effective are existing adversarial attack algorithms on tabular data?
  • RQ2.2: How imperceptible are existing adversarial attacks when evaluated against the defined properties?
  • RQ2.3: To what extent can attacks achieve effectiveness while remaining imperceptible?

Objective 2: Develop an evaluation framework to benchmark adversarial attacks on tabular data.

Research Question & Objective 3

RQ3: How can new adversarial attacks on tabular data be designed to generate both effective and imperceptible adversarial examples?

Sub-questions:

  • RQ3.1: How effectively can new attack methods preserve statistical consistency while minimising perturbations?
  • RQ3.2: How do proposed methods compare with existing attacks in terms of both effectiveness and imperceptibility?
  • RQ3.3: How does adversarial training with newly designed attacks influence model robustness?
  • RQ3.4: How do design choices such as perturbation constraints, hyperparameters, and feature selection strategies affect the attack effectiveness and imperceptibility?

Objective 3: Design new adversarial attack that unify effectiveness with tabular-specific imperceptibility.

Design Science Research (DSR) Methodology

This thesis follows DSR methodology (Peffers et al. 2014; Hevner et al. 2004), guided by the three-cycle view (Hevner 2007):

  • Relevance cycle: practical challenge of adversarial robustness in tabular data
  • Rigor cycle: extends theoretical foundations of adversarial machine learning
  • Design cycle: iterative development & evaluation of artifacts

Artifacts developed:

  • Imperceptibility properties for tabular data
  • Benchmark for systematic evaluation
  • Proposed new attack method for tabular data
Hevner, A. 2007. “The Three Cycle View of Design Science.” Scandinavian Journal of Information Systems 19 (2): 4. https://aisel.aisnet.org/sjis/vol19/iss2/4.
Hevner, March, Park, and Ram. 2004. “Design Science in Information Systems Research.” MIS Quarterly: Management Information Systems 28 (1): 75. https://doi.org/10.2307/25148625.
Peffers, Ken, Tuure Tuunanen, Marcus A Rothenberger, and Samir Chatterjee. 2014. “A Design Science Research Methodology for Information Systems Research.” Journal of Management Information Systems, 45–77. https://doi.org/10.2753/MIS0742-1222240302.

Research Design

Start from common pipeline of adversarial attack on tabular data.

Research Design

RQ1 build the foundation of imperceptibility on tabular data.

Research Design

RQ2 benchmark the adversarial attacks under effectiveness and imperceptibility.

Research Design

RQ3 design the new adversarial attack methods that unify effectiveness with tabular-specific imperceptibility.

Objective 1

🧩 Foundation of Imperceptibility on Tabular Data

Publication:

Zhipeng He, Chun Ouyang, Laith Alzubaidi, Alistair Barros, and Catarina Moreira. “Investigating imperceptibility of adversarial attacks on tabular data: An empirical analysis”. In: Intelligent Systems with Applications 25, 200461 (2025). DOI: 10.1016/j.iswa.2024.200461.

Establishing Criteria and Deriving Properties for Imperceptibility

Quantitative Properties

Proximity

Sparsity

Deviation

Sensitivity
  • Proximity: Measures how close \(\tilde{x}\) is to \(x\).
  • Sparsity: Measures how many features are altered in \(\tilde{x}\).
  • Deviation: Measures how far \(\tilde{x}\) is from the data distribution \(\mathcal{P}(\mathcal{D})\).
  • Sensitivity: Measures the perturbation on narrow-variance features (\(x_0\) and \(x_2\)) over the dataset \(\mathcal{D}\).

Qualitative Properties

Three properties:

  1. Immutability – features that are inherently fixed or should not be altered for ethical or practical reasons.
  2. Feasibility – whether perturbed values remain semantically correct.
  3. Feature Interdependency – how perturbations respect the feature relationships and dependencies.

Illustration of Quantitative Properties
  • Depend on domain knowledge and case-based reasoning.

Foundation for Evaluating Imperceptibility

  • Provides a principled framework to capture imperceptibility in tabular data.
  • Moves beyond ad hoc or isolated criteria used in previous work.

Objective 2

📊 Benchmark of Tabular Attacks

Paper:

Zhipeng He, Chun Ouyang, Lijie Wen, Cong Liu, and Catarina Moreira. “TabAttackBench: A benchmark for adversarial attacks on tabular data”. In: arXiv [cs.LG] (2025). arXiv: 2505.21027.

Submitted to Expert Systems with Applications, under \(1^{st}\) revision.

Framework of Benchmarking

Dataset Selection

Select 10 datasets from tabular benchmarking suite (Gorishniy et al. 2021) under the following criteria:

  1. Classification focus – Each dataset presents classification problems
  2. Feasible size – Feature dimensionality (after one-hot encoding) is capped at ≤ 200
  3. Realistic domains – Representative of real-world domains (health, finance, physics, etc.)

Gorishniy, Yu V, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. 2021. “Revisiting Deep Learning Models for Tabular Data.” In Advances in Neural Information Processing Systems 34 (NeurIPS 2021), edited by Marc’aurelio Ranzato, Alina Beygelzimer, Yann N Dauphin, Percy Liang, and Jennifer Wortman Vaughan, 34:18932–43. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2021/hash/9d86d83f925f2149e9edb0ac3b49229c-Abstract.html.

Predictive Model Selection

  1. Diversity - covering classical and modern architectures
  2. Interpretability - balancing transparency with complexity
  3. Performance - ensuring competitive accuracy on tabular tasks

Adversarial Attack Selection

Evaluation Metrics

Effectiveness

Attack Success Rate (ASR):

  • fraction of perturbed instances that being successfully misclassfied
  • Higher is better

\[ \text{ASR} = \frac{1}{n}\sum_{i=1}^{n}\one( \boldsymbol{x}^{adv}\neq y) \]

Evaluation Metrics

Imperceptibility

4 Quantitative Properties: Lower is better

  • Proximity: Average \(\ell_2\) distance between \(x\) and \(\tilde{x}\).
  • Sparsity: Fraction of features changed in \(\tilde{x}\).
  • Deviation: Outlier rate against the data distribution.
  • Sensitivity: Variance-normalised perturbation size on narrow-variance features.

Imperceptibility Score (IS): Equal-weighted harmonic mean of above normalised properties

\[ \text{IS} = \frac{4}{\frac{1}{\text{Proximity}} + \frac{1}{\text{Sparsity}} + \frac{1}{\text{Deviation}} + \frac{1}{\text{Sensitivity}}} \]

👉 Harmonic mean penalises imbalance (one bad metric drags IS down).

Trade-off analysis: ASR vs Imperceptibility (IS)

ASR vs IS 2D density. Higher attack success rate is better. Lower imperceptibility score is better.

Trade-off analysis: ASR vs Imperceptibility (IS)

Choose the threshold of maximum ASR value (0.659) and the minimum IS value (0.181) of Gaussian Noise.

Trade-off analysis: ASR vs Imperceptibility (IS)

Divided into four sectors by the thresholds.

By using the two thresholds, we can divide the 2D density into four sectors.

Trade-off analysis: ASR vs Imperceptibility (IS)

Guassion noise is ineffective and perceptible.

Trade-off analysis: ASR vs Imperceptibility (IS)

Only DeepFool and C&W lies in the effective & imperceptible sector.

Imperceptibility Insights

  • Sparsity
    • FGSM, PGD and BIM attacks alter nearly all numerical features
    • C&W and DeepFool attacks more selective on numerical features
    • However, categorical features mostly ignored by all attacks
  • Proximity
    • FGSM, PGD and BIM attacks effective but with large perturbations.
    • C&W & DeepFool generate closest examples.
  • Deviation
    • FGSM, PGD and BIM attacks often push samples out of distribution.
    • C&W and DeepFool attacks better preserve data alignment.

The t-SNE visualisation shows that FGSM attack keeps generating out-of-distribution examples on Adult dataset.
  • Sensitivity
    • FGSM, PGD and BIM attacks show high variance and instability.
    • C&W perturbs narrow-range features least.

Imperceptibility Insights

Overall Insight

  • FGSM, PGD and BIM attacks are strong but unrealistic; C&W and DeepFool attacks achieve better balance between effectiveness and imperceptibility.
    👉 Unbounded attacks are the better choice for tabular data.
  • All five attacks fail to handle categorical features during the generation of adversarial examples.
    👉 Further Question: What is the root cause of this issue?

Limitation in Categorical Feature Encoding

  • While one-hot encoding simplifies the handling of categorical features by making them compatible with standard distance measurements (such as \(\ell_p\) norms) used for continuous features, it can introduce more sparse feature space.

Changing one categorical feature requires perturbation on at least two encoded features.

How to Design New Attacks for Tabular Data?

💡 Insight 1: Propose based on unbounded attacks (C&W, DeepFool, etc.).

🧩 Insight 2: Consider the imperceptibility properties in the attack design.

⚙️ Insight 3: Unify both categorical and numerical feature spaces in attacks.

Question: Can we incorporate the above insights to design a new attack for tabular data?

How to Design New Attacks for Tabular Data?

Objective 3

⚙️ Generative VAE-based Attack

Paper:

Zhipeng He, Alexander Stevens, Chun Ouyang, Johannes De Smedt, Alistair Barros, and Catarina Moreira. “Crafting imperceptible on-manifold adversarial attacks for tabular data”. In: arXiv [cs.LG] (2025). arXiv: 2507.10998.

Submitted to Applied Soft Computing, under \(1^{st}\) revision.

Manifold Learning

Manifold Hypothesis: High-dimensional data often lies on a lower-dimensional manifold embedded within the high-dimensional input space.

The Latent Space of Variational Autoencoder (VAE) (Kingma and Welling 2014). Image Source: Tim von Hahn

Unify both categorical and numerical feature spaces into a continuous latent space.

👉 Addressing the issue of categorical feature encoding.

Learn the distribution of original data to generate in-distribution adversarial examples.

👉 Addressing the deviation property of imperceptibility.

Kingma, Diederik P, and Max Welling. 2014. “Auto-Encoding Variational Bayes.” In 2nd International Conference on Learning Representations (ICLR 2014), edited by Yoshua Bengio and Yann LeCun.

Overview of VAE-Based Attack

Step 1: VAE Training:

  1. Optimise \(\mathcal{L}_{\mathrm{VAE}}\) until convergence
  2. Freeze encoder & decoder for adversarial generation

\[ \begin{align} \mathcal{L}_{\mathrm{VAE}} &= \underbrace{ \mathbb{E}_{q_\phi(z\mid x)} \Big[ \underbrace{\|x^{\mathrm{num}} - \hat{x}^{\mathrm{num}}\|^2}_{\substack{\text{Numerical reconstruction} \\ \text{(MSE, Gaussian likelihood)}}} + \underbrace{\mathcal{L}_{\mathrm{cat}}(x^{\mathrm{cat}}, \hat{x}^{\mathrm{cat}})}_{\substack{\text{Categorical reconstruction} \\ \text{(Cross-entropy, Softmax probability)}}} \Big] }_{\text{Reconstruction Loss}} \\ &\quad + \beta \cdot \underbrace{D_{\mathrm{KL}}\big(q_\phi(z\mid x) \parallel p(z)\big)}_{\substack{\text{KL divergence} \\ \text{(Latent regularisation)}}} \,+\, \alpha \cdot \underbrace{\mathcal{L}_{\text{cls}}(h_\omega(z), y)}_{\substack{\text{Classification loss} \\ \text{(Class separation)}}}. \end{align} \]

VAE Architecture

Overview of VAE-Based Attack

Step 2: Adversarial example generation:

  1. Encode input: \(x \to z\)
  2. Add perturbation \(\delta\) in latent space
  3. Decode \((z+\delta) \to \tilde{x}\)
  4. Optimise \(\delta\) via CW-style loss: \(\min_\delta \; \lambda \,\mathcal{L}_{cls}(\tilde{x}, y) + \|\delta\|_2\)

Perturbation in latent space

Evaluation Setup

Better Standardlisation

Z-score normalisation to address the sensitivity property of imperceptibility.

Baseline Manifold Attacks from Computer Vision

  1. PGD-VAE (Stutz, Hein, and Schiele 2019): Applied PGD in latent space of VAE.
  2. Delta-Z (Creswell, Bharath, and Sengupta 2017): perturbs latent vectors multiplicatively to flip class-discriminative feature signs, making it suitable only for binary classification tasks.

New Metric for In-Distribution Success

Proposed In-Distribution Success Rate (IDSR) to measure the success rate of adversarial examples in the in-distribution, i.e., \(\mathrm{IDSR} = \mathrm{ASR} \times (1 - \mathrm{OR})\).

Creswell, Antonia, Anil A Bharath, and Biswa Sengupta. 2017. LatentPoison - Adversarial Attacks on the Latent Space.” arXiv [Cs.LG]. http://arxiv.org/abs/1711.02879.
Stutz, David, Matthias Hein, and Bernt Schiele. 2019. “Disentangling Adversarial Robustness and Generalization.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), 6969–80. IEEE. https://doi.org/10.1109/cvpr.2019.00714.

Reconstruction Evaluation

  1. The trained VAE achieves strong reconstruction performance across all datasets.
  2. The German dataset performs slightly worse, likely due to its limited number of samples.

Reconstruction Evaluation

💡 Introducing the classification-aware loss (\(\mathcal{L}_{cls}\)) encourages the latent space to become more discriminative.

Left t-SNE: Baseline VAE (without \(\mathcal{L}_{cls}\)) on Phishing dataset
classes overlap in latent space.

Right t-SNE: Proposed VAE (with \(\mathcal{L}_{cls}\)) on Phishing dataset
clearer class separation.

Why Use VAE Instead of GAN for Tabular Data?

The VAE’s probabilistic latent structure better preserves both numerical continuity and categorical consistency than Generative Adversarial Network (GAN) (Goodfellow et al. 2014), making it more suitable for imperceptible adversarial example generation on tabular data.

Goodfellow, Ian J, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. “Generative Adversarial Nets.” In Advances in Neural Information Processing Systems 27 (NeurIPS 2014). https://proceedings.neurips.cc/paper_files/paper/2014/hash/f033ed80deb0234979a61f95710dbe25-Abstract.html.

Evaluating Proposed Attack (IDSR + OR)

💡 Our VAE-based attack achieves the best overall In-Distribution Success Rate (IDSR) while maintaining a low Outlier Rate (OR), outperforming other VAE-based methods.

Evaluating Attack Imperceptibility (t-SNE)

💡 The latent space visualisations (t-SNE) show that adversarial examples (green) generated by our VAE-based attack remain closely aligned with original inputs (blue).

Phishing dataset and MLP model.

PenDigits dataset and TabTransformer model.

Applying Our Attack to Adversarial Training

A common practice in adversarial training is to generate adversarial examples during training (Qian et al. 2022).
Qian, Zhuang, Kaizhu Huang, Qiu-Feng Wang, and Xu-Yao Zhang. 2022. “A Survey of Robust Adversarial Training in Pattern Recognition: Fundamental, Theory, and Methodologies.” Pattern Recognition 131. https://doi.org/10.1016/j.patcog.2022.108889.

Applying Our Attack to Adversarial Training

💡 Adversarial training effectively reduces Attack Success Rate (robustness gain) while maintaining relatively high model accuracy across all models.

🌟 Conclusion & Future Directions

Summary

  • Research Problem:

    • How can adversarial examples in tabular domains be both effective and imperceptible?
  • Approach: Three-part progression:

    1. Formalised 7 properties of imperceptibility on tabular data.
    2. Evaluated existing attacks → exposed gaps and opportunities.
    3. Proposed VAE-based on-manifold attack enforcing imperceptibility constraints.
  • Outcome:

    • Framework + benchmark + method = foundations, diagnostics, and artefact for advancing AML in tabular data.

Contributions

Conceptual

  • Property-based framework: statistical consistency, semantic validity, structural integrity.

Analytical

  • Quantitative + qualitative metrics → benchmark exposing drawbacks of existing attacks.

Methodological

  • Generative VAE-based attack unifying categorical & numerical features in latent space.

Practical

  • Realistic adversarial examples → pathway to adversarial training for more robust tabular models.

From theory to tools, the thesis bridges conceptual clarity → systematic evaluation → practical innovation.

Limitations & Future Work

Limitations

  • Benchmark mainly quantitative (qualitative constraints case-based only).
  • Finite datasets, models, attack families.
  • On-manifold attack depends on sufficient training data quality.

Future Work

  • LLMs for imperceptibility → encode semantic/feasibility constraints.
  • Non-uniform perturbations → adaptive perturbation by feature importance & semantics.
  • Beyond white-box → explore grey-box & black-box attack settings.

Closing Insight

Effort to “mend the fence before the sheep are lost”:
→ strengthening adversarial robustness for tabular data models.

Adversarial attacks misclassify models via “broken fence”.

Proactive adversarial training “mending the fence”.

References

Ahmed, Murtadha, Bo Wen, Luo Ao, and Yunfeng Liu. 2025. “Integrated Gradients-Based Defense Against Adversarial Word Substitution Attacks.” Neural Computing & Applications 37 (18): 12921–39. https://doi.org/10.1007/s00521-025-11135-3.
Assion, Felix, Peter Schlicht, Florens Gressner, Wiebke Gunther, Fabian Huger, Nico Schmidt, and Umair Rasheed. 2019. “The Attack Generator: A Systematic Approach Towards Constructing Adversarial Attacks.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW 2019), 1370–79. IEEE. https://doi.org/10.1109/cvprw.2019.00177.
Ballet, Vincent, Xavier Renard, Jonathan Aigrain, Thibault Laugel, Pascal Frossard, and Marcin Detyniecki. 2019. “Imperceptible Adversarial Attacks on Tabular Data.” arXiv [Stat.ML]. http://arxiv.org/abs/1911.03274.
Biggio, Battista, and Fabio Roli. 2018. “Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning.” Pattern Recognition 84: 317–31. https://doi.org/10.1016/j.patcog.2018.07.023.
Borisov, Vadim, Tobias Leemann, Kathrin Sebler, Johannes Haug, Martin Pawelczyk, and Gjergji Kasneci. 2024. “Deep Neural Networks and Tabular Data: A Survey.” IEEE Transactions on Neural Networks and Learning Systems 35 (6): 7499–519. https://doi.org/10.1109/TNNLS.2022.3229161.
Carlini, Nicholas, and David Wagner. 2017. “Towards Evaluating the Robustness of Neural Networks.” In 2017 IEEE Symposium on Security and Privacy (SP 2017), 39–57. IEEE. https://doi.org/10.1109/sp.2017.49.
———. 2018. “Audio Adversarial Examples: Targeted Attacks on Speech-to-Text.” In 2018 IEEE Security and Privacy Workshops (SP Workshop 2018), 1–7. IEEE. https://doi.org/10.1109/spw.2018.00009.
Chernikova, Alesia, and Alina Oprea. 2022. FENCE: Feasible Evasion Attacks on Neural Networks in Constrained Environments.” ACM Transactions on Privacy and Security 25 (4): 1–34. https://doi.org/10.1145/3544746.
Cinà, Antonio Emanuele, Jérôme Rony, Maura Pintor, Luca Demetrio, Ambra Demontis, Battista Biggio, Ismail Ben Ayed, and Fabio Roli. 2025. AttackBench: Evaluating Gradient-Based Attacks for Adversarial Examples.” In Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence (AAAI-25), edited by Toby Walsh, Julie Shah, and Zico Kolter, 2600–2608. Association for the Advancement of Artificial Intelligence (AAAI). https://doi.org/10.1609/aaai.v39i3.32263.
Creswell, Antonia, Anil A Bharath, and Biswa Sengupta. 2017. LatentPoison - Adversarial Attacks on the Latent Space.” arXiv [Cs.LG]. http://arxiv.org/abs/1711.02879.
Croce, Francesco, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. 2021. RobustBench: A Standardized Adversarial Robustness Benchmark.” In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS Datasets and Benchmarks 2021), edited by Joaquin Vanschoren and Sai-Kit Yeung. Vol. 1. Curran Associates, Inc. https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/a3c65c2974270fd093ee8a9bf8ae7d0b-Abstract-round2.html.
Goodfellow, Ian J, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. “Generative Adversarial Nets.” In Advances in Neural Information Processing Systems 27 (NeurIPS 2014). https://proceedings.neurips.cc/paper_files/paper/2014/hash/f033ed80deb0234979a61f95710dbe25-Abstract.html.
Goodfellow, Ian J, Jonathon Shlens, and Christian Szegedy. 2015. “Explaining and Harnessing Adversarial Examples.” In 3rd International Conference on Learning Representations (ICLR 2015), edited by Yoshua Bengio and Yann LeCun. https://arxiv.org/abs/1412.6572.
Gorishniy, Yu V, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. 2021. “Revisiting Deep Learning Models for Tabular Data.” In Advances in Neural Information Processing Systems 34 (NeurIPS 2021), edited by Marc’aurelio Ranzato, Alina Beygelzimer, Yann N Dauphin, Percy Liang, and Jennifer Wortman Vaughan, 34:18932–43. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2021/hash/9d86d83f925f2149e9edb0ac3b49229c-Abstract.html.
Hevner, A. 2007. “The Three Cycle View of Design Science.” Scandinavian Journal of Information Systems 19 (2): 4. https://aisel.aisnet.org/sjis/vol19/iss2/4.
Hevner, March, Park, and Ram. 2004. “Design Science in Information Systems Research.” MIS Quarterly: Management Information Systems 28 (1): 75. https://doi.org/10.2307/25148625.
Hingun, Nabeel, Chawin Sitawarin, Jerry Li, and David Wagner. 2023. REAP: A Large-Scale Realistic Adversarial Patch Benchmark.” In 2023 IEEE/CVF International Conference on Computer Vision (ICCV 2023), 4640–51. IEEE. https://doi.org/10.1109/ICCV51070.2023.00428.
IEEE. 2008. IEEE Standard Computer Dictionary: A Compilation of IEEE Standard Computer Glossaries.” Piscataway, NJ, USA: IEEE. https://doi.org/10.1109/ieeestd.1991.106963.
Kingma, Diederik P, and Max Welling. 2014. “Auto-Encoding Variational Bayes.” In 2nd International Conference on Learning Representations (ICLR 2014), edited by Yoshua Bengio and Yann LeCun.
Kurakin, Alexey, Ian Goodfellow, and Samy Bengio. 2016. “Adversarial Examples in the Physical World.” arXiv [Cs.CV]. http://arxiv.org/abs/1607.02533.
Li, Bo, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou. 2023. “Trustworthy AI: From Principles to Practices.” ACM Computing Surveys 55 (9): 1–46. https://doi.org/10.1145/3555803.
Li, Deqiang, Qianmu Li, Yanfang (fanny) Ye, and Shouhuai Xu. 2023. “Arms Race in Adversarial Malware Detection: A Survey.” ACM Computing Surveys 55 (1): 1–35. https://doi.org/10.1145/3484491.
Lunghi, Daniele, Alkis Simitsis, Olivier Caelen, and Gianluca Bontempi. 2023. “Adversarial Learning in Real-World Fraud Detection: Challenges and Perspectives.” In Proceedings of the Second ACM Data Economy Workshop (DEC 2023), 27–33. New York, NY, USA: ACM. https://doi.org/10.1145/3600046.3600051.
Madry, Aleksander, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. “Towards Deep Learning Models Resistant to Adversarial Attacks.” In 6th International Conference on Learning Representations (ICLR 2018). OpenReview.net. https://openreview.net/forum?id=rJzIBfZAb.
Moosavi-Dezfooli, Seyed-Mohsen, Alhussein Fawzi, and Pascal Frossard. 2016. DeepFool: A Simple and Accurate Method to Fool Deep Neural Networks.” In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), 2574–82. IEEE. https://doi.org/10.1109/cvpr.2016.282.
Papernot, Nicolas, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami. 2016. “The Limitations of Deep Learning in Adversarial Settings.” In 2016 IEEE European Symposium on Security and Privacy (EuroS&p 2016), 372–87. IEEE. https://doi.org/10.1109/eurosp.2016.36.
Peffers, Ken, Tuure Tuunanen, Marcus A Rothenberger, and Samir Chatterjee. 2014. “A Design Science Research Methodology for Information Systems Research.” Journal of Management Information Systems, 45–77. https://doi.org/10.2753/MIS0742-1222240302.
Pelekis, Sotiris, Thanos Koutroubas, Afroditi Blika, Anastasis Berdelis, Evangelos Karakolis, Christos Ntanos, Evangelos Spiliotis, and Dimitris Askounis. 2025. “Adversarial Machine Learning: A Review of Methods, Tools, and Critical Industry Sectors.” Artificial Intelligence Review 58 (8): 1–87. https://doi.org/10.1007/s10462-025-11147-4.
Qian, Zhuang, Kaizhu Huang, Qiu-Feng Wang, and Xu-Yao Zhang. 2022. “A Survey of Robust Adversarial Training in Pattern Recognition: Fundamental, Theory, and Methodologies.” Pattern Recognition 131. https://doi.org/10.1016/j.patcog.2022.108889.
Siddiqui, Shoaib Ahmed, Andreas Dengel, and Sheraz Ahmed. 2020. “Benchmarking Adversarial Attacks and Defenses for Time-Series Data.” In 27th International Conference on Neural Information Processing (ICONIP 2020), 12534:544–54. Lecture Notes in Computer Science. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-63836-8_45.
Simonetto, Thibault, Salah Ghamizi, and Maxime Cordy. 2024. “Towards Adaptive Attacks on Constrained Tabular Machine Learning.” In ICML 2024 Workshop on the Next Generation of AI Safety. OpenReview.net. https://openreview.net/forum?id=DnvYdmR9OB.
Stutz, David, Matthias Hein, and Bernt Schiele. 2019. “Disentangling Adversarial Robustness and Generalization.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), 6969–80. IEEE. https://doi.org/10.1109/cvpr.2019.00714.
Sun, Mengying, Fengyi Tang, Jinfeng Yi, Fei Wang, and Jiayu Zhou. 2018. “Identify Susceptible Locations in Medical Records via Adversarial Attacks on Deep Predictive Models.” In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD 2018), edited by Yike Guo and Faisal Farooq. New York, NY, USA: ACM. https://doi.org/10.1145/3219819.3219909.
Szegedy, Christian, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2014. “Intriguing Properties of Neural Networks.” In 2nd International Conference on Learning Representations (ICLR 2014), edited by Yoshua Bengio and Yann LeCun. https://arxiv.org/abs/1312.6199.
Zheng, Qinkai, Xu Zou, Yuxiao Dong, Yukuo Cen, Da Yin, Jiarong Xu, Yang Yang, and Jie Tang. 2021. “Graph Robustness Benchmark: Benchmarking the Adversarial Robustness of Graph Machine Learning.” In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS Datasets and Benchmarks 2021), edited by Joaquin Vanschoren and Sai-Kit Yeung. Curran Associates, Inc. https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/6cdd60ea0045eb7a6ec44c54d29ed402-Abstract-round2.html.
Ahmed, Murtadha, Bo Wen, Luo Ao, and Yunfeng Liu. 2025. “Integrated Gradients-Based Defense Against Adversarial Word Substitution Attacks.” Neural Computing & Applications 37 (18): 12921–39. https://doi.org/10.1007/s00521-025-11135-3.
Assion, Felix, Peter Schlicht, Florens Gressner, Wiebke Gunther, Fabian Huger, Nico Schmidt, and Umair Rasheed. 2019. “The Attack Generator: A Systematic Approach Towards Constructing Adversarial Attacks.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW 2019), 1370–79. IEEE. https://doi.org/10.1109/cvprw.2019.00177.
Ballet, Vincent, Xavier Renard, Jonathan Aigrain, Thibault Laugel, Pascal Frossard, and Marcin Detyniecki. 2019. “Imperceptible Adversarial Attacks on Tabular Data.” arXiv [Stat.ML]. http://arxiv.org/abs/1911.03274.
Biggio, Battista, and Fabio Roli. 2018. “Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning.” Pattern Recognition 84: 317–31. https://doi.org/10.1016/j.patcog.2018.07.023.
Borisov, Vadim, Tobias Leemann, Kathrin Sebler, Johannes Haug, Martin Pawelczyk, and Gjergji Kasneci. 2024. “Deep Neural Networks and Tabular Data: A Survey.” IEEE Transactions on Neural Networks and Learning Systems 35 (6): 7499–519. https://doi.org/10.1109/TNNLS.2022.3229161.
Carlini, Nicholas, and David Wagner. 2017. “Towards Evaluating the Robustness of Neural Networks.” In 2017 IEEE Symposium on Security and Privacy (SP 2017), 39–57. IEEE. https://doi.org/10.1109/sp.2017.49.
———. 2018. “Audio Adversarial Examples: Targeted Attacks on Speech-to-Text.” In 2018 IEEE Security and Privacy Workshops (SP Workshop 2018), 1–7. IEEE. https://doi.org/10.1109/spw.2018.00009.
Chernikova, Alesia, and Alina Oprea. 2022. FENCE: Feasible Evasion Attacks on Neural Networks in Constrained Environments.” ACM Transactions on Privacy and Security 25 (4): 1–34. https://doi.org/10.1145/3544746.
Cinà, Antonio Emanuele, Jérôme Rony, Maura Pintor, Luca Demetrio, Ambra Demontis, Battista Biggio, Ismail Ben Ayed, and Fabio Roli. 2025. AttackBench: Evaluating Gradient-Based Attacks for Adversarial Examples.” In Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence (AAAI-25), edited by Toby Walsh, Julie Shah, and Zico Kolter, 2600–2608. Association for the Advancement of Artificial Intelligence (AAAI). https://doi.org/10.1609/aaai.v39i3.32263.
Creswell, Antonia, Anil A Bharath, and Biswa Sengupta. 2017. LatentPoison - Adversarial Attacks on the Latent Space.” arXiv [Cs.LG]. http://arxiv.org/abs/1711.02879.
Croce, Francesco, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. 2021. RobustBench: A Standardized Adversarial Robustness Benchmark.” In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS Datasets and Benchmarks 2021), edited by Joaquin Vanschoren and Sai-Kit Yeung. Vol. 1. Curran Associates, Inc. https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/a3c65c2974270fd093ee8a9bf8ae7d0b-Abstract-round2.html.
Goodfellow, Ian J, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. “Generative Adversarial Nets.” In Advances in Neural Information Processing Systems 27 (NeurIPS 2014). https://proceedings.neurips.cc/paper_files/paper/2014/hash/f033ed80deb0234979a61f95710dbe25-Abstract.html.
Goodfellow, Ian J, Jonathon Shlens, and Christian Szegedy. 2015. “Explaining and Harnessing Adversarial Examples.” In 3rd International Conference on Learning Representations (ICLR 2015), edited by Yoshua Bengio and Yann LeCun. https://arxiv.org/abs/1412.6572.
Gorishniy, Yu V, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. 2021. “Revisiting Deep Learning Models for Tabular Data.” In Advances in Neural Information Processing Systems 34 (NeurIPS 2021), edited by Marc’aurelio Ranzato, Alina Beygelzimer, Yann N Dauphin, Percy Liang, and Jennifer Wortman Vaughan, 34:18932–43. Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2021/hash/9d86d83f925f2149e9edb0ac3b49229c-Abstract.html.
Hevner, A. 2007. “The Three Cycle View of Design Science.” Scandinavian Journal of Information Systems 19 (2): 4. https://aisel.aisnet.org/sjis/vol19/iss2/4.
Hevner, March, Park, and Ram. 2004. “Design Science in Information Systems Research.” MIS Quarterly: Management Information Systems 28 (1): 75. https://doi.org/10.2307/25148625.
Hingun, Nabeel, Chawin Sitawarin, Jerry Li, and David Wagner. 2023. REAP: A Large-Scale Realistic Adversarial Patch Benchmark.” In 2023 IEEE/CVF International Conference on Computer Vision (ICCV 2023), 4640–51. IEEE. https://doi.org/10.1109/ICCV51070.2023.00428.
IEEE. 2008. IEEE Standard Computer Dictionary: A Compilation of IEEE Standard Computer Glossaries.” Piscataway, NJ, USA: IEEE. https://doi.org/10.1109/ieeestd.1991.106963.
Kingma, Diederik P, and Max Welling. 2014. “Auto-Encoding Variational Bayes.” In 2nd International Conference on Learning Representations (ICLR 2014), edited by Yoshua Bengio and Yann LeCun.
Kurakin, Alexey, Ian Goodfellow, and Samy Bengio. 2016. “Adversarial Examples in the Physical World.” arXiv [Cs.CV]. http://arxiv.org/abs/1607.02533.
Li, Bo, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou. 2023. “Trustworthy AI: From Principles to Practices.” ACM Computing Surveys 55 (9): 1–46. https://doi.org/10.1145/3555803.
Li, Deqiang, Qianmu Li, Yanfang (fanny) Ye, and Shouhuai Xu. 2023. “Arms Race in Adversarial Malware Detection: A Survey.” ACM Computing Surveys 55 (1): 1–35. https://doi.org/10.1145/3484491.
Lunghi, Daniele, Alkis Simitsis, Olivier Caelen, and Gianluca Bontempi. 2023. “Adversarial Learning in Real-World Fraud Detection: Challenges and Perspectives.” In Proceedings of the Second ACM Data Economy Workshop (DEC 2023), 27–33. New York, NY, USA: ACM. https://doi.org/10.1145/3600046.3600051.
Madry, Aleksander, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. “Towards Deep Learning Models Resistant to Adversarial Attacks.” In 6th International Conference on Learning Representations (ICLR 2018). OpenReview.net. https://openreview.net/forum?id=rJzIBfZAb.
Moosavi-Dezfooli, Seyed-Mohsen, Alhussein Fawzi, and Pascal Frossard. 2016. DeepFool: A Simple and Accurate Method to Fool Deep Neural Networks.” In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), 2574–82. IEEE. https://doi.org/10.1109/cvpr.2016.282.
Papernot, Nicolas, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami. 2016. “The Limitations of Deep Learning in Adversarial Settings.” In 2016 IEEE European Symposium on Security and Privacy (EuroS&p 2016), 372–87. IEEE. https://doi.org/10.1109/eurosp.2016.36.
Peffers, Ken, Tuure Tuunanen, Marcus A Rothenberger, and Samir Chatterjee. 2014. “A Design Science Research Methodology for Information Systems Research.” Journal of Management Information Systems, 45–77. https://doi.org/10.2753/MIS0742-1222240302.
Pelekis, Sotiris, Thanos Koutroubas, Afroditi Blika, Anastasis Berdelis, Evangelos Karakolis, Christos Ntanos, Evangelos Spiliotis, and Dimitris Askounis. 2025. “Adversarial Machine Learning: A Review of Methods, Tools, and Critical Industry Sectors.” Artificial Intelligence Review 58 (8): 1–87. https://doi.org/10.1007/s10462-025-11147-4.
Qian, Zhuang, Kaizhu Huang, Qiu-Feng Wang, and Xu-Yao Zhang. 2022. “A Survey of Robust Adversarial Training in Pattern Recognition: Fundamental, Theory, and Methodologies.” Pattern Recognition 131. https://doi.org/10.1016/j.patcog.2022.108889.
Siddiqui, Shoaib Ahmed, Andreas Dengel, and Sheraz Ahmed. 2020. “Benchmarking Adversarial Attacks and Defenses for Time-Series Data.” In 27th International Conference on Neural Information Processing (ICONIP 2020), 12534:544–54. Lecture Notes in Computer Science. Cham: Springer International Publishing. https://doi.org/10.1007/978-3-030-63836-8_45.
Simonetto, Thibault, Salah Ghamizi, and Maxime Cordy. 2024. “Towards Adaptive Attacks on Constrained Tabular Machine Learning.” In ICML 2024 Workshop on the Next Generation of AI Safety. OpenReview.net. https://openreview.net/forum?id=DnvYdmR9OB.
Stutz, David, Matthias Hein, and Bernt Schiele. 2019. “Disentangling Adversarial Robustness and Generalization.” In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), 6969–80. IEEE. https://doi.org/10.1109/cvpr.2019.00714.
Sun, Mengying, Fengyi Tang, Jinfeng Yi, Fei Wang, and Jiayu Zhou. 2018. “Identify Susceptible Locations in Medical Records via Adversarial Attacks on Deep Predictive Models.” In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD 2018), edited by Yike Guo and Faisal Farooq. New York, NY, USA: ACM. https://doi.org/10.1145/3219819.3219909.
Szegedy, Christian, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2014. “Intriguing Properties of Neural Networks.” In 2nd International Conference on Learning Representations (ICLR 2014), edited by Yoshua Bengio and Yann LeCun. https://arxiv.org/abs/1312.6199.
Zheng, Qinkai, Xu Zou, Yuxiao Dong, Yukuo Cen, Da Yin, Jiarong Xu, Yang Yang, and Jie Tang. 2021. “Graph Robustness Benchmark: Benchmarking the Adversarial Robustness of Graph Machine Learning.” In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS Datasets and Benchmarks 2021), edited by Joaquin Vanschoren and Sai-Kit Yeung. Curran Associates, Inc. https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/6cdd60ea0045eb7a6ec44c54d29ed402-Abstract-round2.html.

Thank you!

  • XAMI LAB
  • zhipeng.he@hdr.qut.edu.au
  • github.com/ZhipengHe
  • zhipenghe.me

Check out the slides here 👇

🔗 https://slides.zhipenghe.me/2025-PhD-Final-Seminar