ISCSR Research Publishing
Journal of Adaptive Data and Algorithmic Intelligence

Agentic SQL Generation for Reliable Data Warehouse Query Automation

Read & download PDF
Abstract

Natural-language query generation has become important for enterprise data warehouses, but generated SQL often contains incorrect joins, missing filters, aggregation errors, inefficient execution plans, and schema hallucinations. This study investigates reliable SQL generation through agentic query validation and optimization. We propose SQLAgent-Opt, a multi-agent model consisting of a schema interpretation agent, a query drafting agent, an execution validation agent, a semantic checking agent, and a cost optimization agent. The schema interpretation agent extracts table relationships, primary keys, foreign keys, column types, and business constraints from database metadata. The query drafting agent converts natural-language questions into executable SQL. The execution validation agent checks syntax, runtime errors, and empty-result anomalies. The semantic checking agent verifies whether selected columns, joins, grouping rules, and filtering conditions match the original question. The cost optimization agent rewrites queries using index-aware predicates, join reordering, subquery simplification, and aggregation pruning. Experiments were conducted on 1,540 natural-language query tasks from finance, healthcare, retail, and logistics data warehouses, covering 326 relational tables, 4,870 columns, and 18.6 million records. Compared with a single-agent SQL generator, SQLAgent-Opt improved exact-match accuracy from 58.9% to 73.6% and execution accuracy from 66.4% to 82.1%. The rate of incorrect join paths decreased from 19.7% to 7.8%, while aggregation-related errors decreased by 34.5%. On large tables containing more than one million records, the optimized queries reduced average execution time by 41.2% and temporary memory usage by 26.8%. These results show that agentic SQL generation can improve both semantic correctness and execution efficiency in data warehouse automation.

Keywords
SQL generationdata warehousequery optimizationmulti-agent systemssemantic validationnatural-language interfacedatabase automation
References
  1. Jiao, Y., Shi, T., Zhao, B., & Wang, A. (2026). Retrieval-Guided Structured Reasoning and Interpretable Representation Learning for Large-Scale Video–Language Models.
  2. Baig, M. S., Sher, T., Rehman, A., & Sheikh, S. (2025). A systematic literature review of text-to-sql: Performance, challenges, and limitations. ICCK Transactions on Advanced Computing and Systems, 2(1), 1-24.
  3. Liang, S., Hsu, K. H., & Liu, Z. (2026). Rolling Scheduling and Capacity Balancing for Server Assembly Under Engineering Change Disturbances.
  4. Zheng, J., & Makar, M. (2022). Causally motivated multi-shortcut identification and removal. Advances in Neural Information Processing Systems, 35, 12800-12812.
  5. Fan, Y., Miyazaki, T., Tang, Z., Wang, J., Huang, Y., & Omachi, S. (2025, November). Scene Text Reconstructor: A Contextual-Aware Masking Framework for Pre-training Text Detectors. In International Conference on Neural Information Processing (pp. 181-195). Singapore: Springer Nature Singapore.
  6. Yuan, Y., Xu, T., Yin, J., & Huang, J. (2026). Modeling Loading Anomalies and Identifying Root Causes in Media Consumption Workflows on Large Social Media Platforms. Available at SSRN 7393318.
  7. Wang, T., & Xia, Z. (2025). Stability of In-Context Learning: A Spectral Coverage Perspective. arXiv preprint arXiv:2509.20677.
  8. Zhang, G., Xu, Y., Zhang, M., Wang, S., & Lin, N. (2021). The brain network in support of social semantic accumulation. Social cognitive and affective neuroscience, 16(4), 393-405.
  9. Bao, Y., Qiu, Y., & Wang, H. (2026). FLARE: A Real-Time Framework for Detecting Malicious Account Activities at Internet Scale. Available at SSRN 7185659.
  10. Mini Han Wang , " AI-Powered Innovations in Ophthalmic Diagnosis and Treatment ", Bentham Science Publishers (2025). https://doi.org/10.2174/97988988122871250101.
  11. Wang, S., Feng, Y., & Fang, X. (2026, May). A Large Language Model-Enabled Multi-Agent Collaboration Method for Complex Task Solving. In 2026 6th International Symposium on Computer Technology and Information Science (ISCTIS) (pp. 253-256). IEEE.
  12. Yang, Q., & Du, Y. (2026). Predicting Rollback Risks and Controlling Gradual Rollout Traffic for Online Ranking Feature Deployments.
  13. Zhang, Z. (2026). A Study on the Identification of Manipulative Design in Subscription and Payment Interfaces of Digital Consumer Platforms and Its Behavioral Effects. Available at SSRN 6734760.
  14. Liu, W., Zhao, R., Sha, Z., Cui, Q., & Zhang, Y. (2026). Mapping the City Through the Lens of Language Models. arXiv preprint arXiv:2608.02971..
  15. Zhang, Z. (2026). A Study on the Impact of Cost Presentation on Decision-making Bias and Cancellation Behavior in Online Subscriptions. Available at SSRN 7349839.
  16. Cui, J., Zhang, G., Chen, Z., & Yu, N. (2022). Multi-homed abnormal behavior detection algorithm based on fuzzy particle swarm cluster in user and entity behavior analytics. Scientific Reports, 12(1), 22349.
  17. Qi, C., & Qiao, X. (2026). Efficient Data Sampling and Feature Selection Algorithms for Scalable Machine Learning Pipelines.
  18. Zhao, R., Liu, W., Sha, Z., Su, N., Zhang, Y., & Long, Y. (2026). Culturally uneven urban perception in large language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.20048v3.
  19. Xiong, W., Zeng, Y., & Guo, Y. (2026). A Study on the Characterization of Building MEP System States and the Learning of Behavioral Patterns Driven by Large-Scale Operational Data. Available at SSRN 7121883.
  20. Hong, Z., Liang, S., Xu, T., & Chen, H. (2026). A Study on Failover Verification and Recovery Objective Prediction for Cross-Region Cloud Services.
  21. Han, Z., Chen, W., Han, Y., Mao, R., & Qin, J. (2026). Fast Diversified Top-k Rule Discovery via User-Guided Embeddings. IEEE Transactions on Knowledge and Data Engineering.
  22. Huang, J., Xu, T., Yang, J., & Yin, J. (2026). A Graph Neural Network Approach for Financial Data Validation and Risk Identification in Large-Scale Intercompany Transactions. Available at SSRN 7175058.
  23. Zhao, Z., & Welsch, R. E. (2024). Hierarchical reinforced trader (hrt): A bi-level approach for optimizing stock selection and execution. arXiv preprint arXiv:2410.14927.
  24. Quan, Q., Gao, Y., & Wang, Q. (2025). The impact of teacher emotional support on students' engagement in AI-mediated English learning environments: The mediating role of resilience and self-efficacy. Acta Psychologica, 260, 105766.
  25. Yang, Z., Alexandrova, A. N., & Sautet, P. (2026). Modeling CO2 Hydrogenation to Methanol on an Ensemble of Inverse ZrO2 on Cu Catalytic Sites: Mechanism, Reactivity, and Deactivation. Angewandte Chemie, e5448247.
  26. Jiang, C. (2026). Information doesn’t work? Testing experiential drivers of VR application in heritage tourism by affective-cognitive pathways. Journal of Hospitality and Tourism Technology, 1-28.
  27. Gu, X. (2026). Identifying Causal Effects and Analyzing Heterogeneity of User Growth Interventions on Digital Platforms: Evidence from Large-Scale Behavioral Data. Available at SSRN 6809181.
  28. Wang, Z., Yu, J., Liu, H., Zheng, Z., Jin, Y., Li, S., ... & Zhang, K. (2025, July). Generative Music Models’ Alignment with Professional and Amateur Users’ Expectations. In Findings of the Association for Computational Linguistics: ACL 2025 (pp. 6909-6920).
  29. Zhao, J., Fan, J., & Li, L. (2026). A Study on an Explainable Causal-Enhanced LLM Agent for Predicting the Forming Quality of Automotive Component Materials.
  30. You, S. (2026). Verifiable Audit Mechanisms in AI Compliance Automation: Scalability. Available at SSRN 6547458.
  31. Xu, Y., Sun, Y., Zhai, B., Li, M., Liang, W., Li, Y., & Du, S. (2025, April). Zero-shot video moment retrieval via off-the-shelf multimodal large language models. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 39, No. 9, pp. 8978-8986).
  32. Yuan, Y., Yin, J., Huang, J., & Xu, T. (2026). Weakly Supervised Anomaly Detection and Privacy Risk Scoring in Community Message Threads.
  33. Wu, Y., & Su, D. (2026). An Evaluation of Parental Question Characteristics and the Quality of AI Recommendations in Family Counseling for Children with Autism.
  34. Su, J., Cai, C., Zhu, F., He, C., Xu, X., Guan, D., & Si, C. (2024, September). Momentum auxiliary network for supervised local learning. In European Conference on Computer Vision (pp. 276-292). Cham: Springer Nature Switzerland.
  35. Zheng, X., Chen, X., Gong, S., Griffin, X., & Slabaugh, G. (2025, September). Xfmamba: Cross-fusion mamba for multi-view medical image classification. In International Conference on Medical Image Computing and Computer-Assisted Intervention (pp. 672-682). Cham: Springer Nature Switzerland.
  36. Su, J., Zhu, F., Shi, H., Han, T., Qiu, Y., Luo, J., ... & Gao, J. (2025). MAN++: Scaling Momentum Auxiliary Network for Supervised Local Learning in Vision Tasks. arXiv preprint arXiv:2507.16279.
  37. Li, P., Zhang, H., Li, W., Huang, D., Guo, Z., Chen, J., ... & Koshizuka, N. (2025). GeoAvatar: A big mobile phone positioning data-driven method for individualized pseudo personal mobility data generation. Computers, Environment and Urban Systems, 119, 102252.
  38. Lyu, Y., & Ding, R. (2025). Discovering Self-Regulated Learning Patterns in Chatbot-Powered Education Environment. arXiv preprint arXiv:2510.01275.
  39. Xi, X., Chen, P., Qiu, Y., Luo, R., Tong, P., & Liang, J. Video-BCI: Bayesian Cognitive Integration of Self-Prior Hypotheses for Video Understanding. In Forty-third International Conference on Machine Learning.
  40. Liang, J., Xiao, Y., Yang, H., Yu, Z., Liu, J. N., & Yang, K. (2026, April). C-HyPOD: Causal Hyperbolic Representation Learning with Prototype Orthogonal Disentanglement for Graph Out-of-Distribution Recommendation. In Proceedings of the ACM Web Conference 2026 (pp. 6541-6550).
  41. Xu, Z., Zhang, X., Li, R., Tang, Z., Huang, Q., & Zhang, J. (2024). Fakeshield: Explainable image forgery detection and localization via multi-modal large language models. arXiv preprint arXiv:2410.02761.
  42. Jia, Z., You, K., He, W., Tian, Y., Feng, Y., Wang, Y., ... & Zhang, Z. (2023). Event-based semantic segmentation with posterior attention. IEEE Transactions on Image Processing, 32, 1829-1842.
  43. C. Huang, X. Chen, G. Chen, P. Xiao, G. Ye Li and W. Huang, "Deep Reinforcement Learning-Based Resource Allocation for Hybrid Bit and Generative Semantic Communications in Space-Air-Ground Integrated Networks," in IEEE Journal on Selected Areas in Communications, vol. 43, no. 12, pp. 3942-3954, Dec. 2025, doi: 10.1109/JSAC.2025.3623157.
Publication details
Journal
Journal of Adaptive Data and Algorithmic Intelligence
Volume
1 (2026)
Issue
1 · Forthcoming issue
Article number
jadai20260011
License
CC BY 4.0