[论文] Quantum Feature Engineering for Credit Default Prediction: When and Why IQP Circuits Help Linear Classifiers (arXiv:2609.10505)

## 论文概要 **研究领域**: ML **作者**: Menachem Finkelstein, Dian...

论文概要

研究领域: ML 作者: Menachem Finkelstein, Diana Legziel Levy, Zohar Yakhini, Sarel Cohen 发布时间: 2026-09-09 arXiv: 2609.10505

中文摘要

信用违约预测是一个表格分类问题,F1分数的微小提升直接转化为金融风险敞口的降低。本文探究瞬时量子多项式时间(IQP)电路能否产生改善分类器性能的特征,超越原始经典基线和核PCA(最强的无监督经典非线性替代方法)。在UCI信用卡违约数据集上使用五折交叉验证,发现将16个IQP特征(n=8量子比特)附加到逻辑回归模型上将F1从0.462提升到0.517(+0.055, p<0.0001)。核PCA在相同特征数下仅达到0.493。没有其他分类器(随机森林、SVM、XGBoost、k-NN)受益,这指向线性可表达性机制而非通用改进。

原文摘要

Credit default prediction is a tabular classification problem in which modest gains in F1 translate directly into reduced financial exposure. We ask whether Instantaneous Quantum Polynomial-time (IQP) circuits can produce features that improve a classifier over both its raw classical baseline and Kernel PCA – the strongest unsupervised classical non-linear alternative – at an equal feature budget. The dataset provides 23 financial attributes per client; for an n-qubit circuit we select n of them, encode each as a rotation angle, and read 2n expectation values back out as new features. The motivation for using a quantum circuit is computational: an n-qubit IQP circuit runs in constant depth and encodes feature correlations in a 2^n-dimensional Hilbert space, whereas classical simulation of its exact output statistics scales exponentially in n. Using the UCI Default of Credit Card Clients dataset and five-fold cross-validation, we find that appending 16 IQP features (n = 8 qubits) to a Logistic Regression model raises F1 from 0.462 to 0.517 (+0.055, p < 0.0001). Kernel PCA, the next-best method, reaches only 0.493 at the same feature count; the gap survives Benjamini-Hochberg correction across 12 tests (p = 0.00007). No other classifier – Random Forest, SVM, XGBoost, or k-NN – benefits, which points to a linear-expressivity mechanism rather than a generic improvement. We also show that how the 8 input features are chosen matters: Random Forest importance-guided selection reaches F1 = 0.523, while encoding maximally uncorrelated features drops it to 0.496, demonstrating that the circuit amplifies informative structure rather than creating it from scratch.

自动采集于 2026-09-11

#论文 #arXiv #ML #小凯

发表回复

人生梦想 - 关注前沿的计算机技术 acejoy.com 🐾 步子哥の博客 🐾 背多分论坛 🐾 借一步网 🐾 智柴网 沪ICP备2024052574号-1