ISCSR Research Publishing
Systems, Networks and Security Challenges

Interaction-Level Test Script Generation for Web and Mobile Applications

Read & download PDF
Abstract

Automated user-interface testing is essential for modern web and mobile applications, but writing stable test scripts is labor-intensive. Generated test scripts often fail because they depend on fragile selectors, ignore asynchronous loading, miss important user paths, or contain weak test oracles. This study examines interaction-level test script generation for web and mobile applications. We propose TestNav-Agent, a multi-agent model that includes a screen-state analysis agent, a user-intent planning agent, a test script generation agent, an oracle construction agent, and a flakiness repair agent. The screen-state analysis agent builds page and component representations from DOM trees, accessibility labels, screenshots, and event traces. The user-intent planning agent converts functional requirements into realistic interaction paths. The test script generation agent produces Playwright, Selenium, and Appium scripts. The oracle construction agent checks expected page transitions, data changes, form validation messages, and error states. The flakiness repair agent replaces unstable selectors, adds wait conditions, and removes timing-sensitive assertions. The study was conducted on 52 production web and mobile interfaces from e-commerce, online education, healthcare appointment, and financial service applications. The evaluation material included 68,000 recorded user actions, 9,400 screen states, 1,260 bug reports, and 740 manually written regression tests. TestNav-Agent generated scripts that covered 76.8% of critical user paths identified by product teams, compared with 54.2% for a direct script-generation model. Across 100 repeated test runs, the flaky failure rate decreased from 13.5% to 4.1%. The generated tests detected 37 previously documented regression bugs and 11 unreleased interface defects during staging evaluation. Selector repair was especially effective in dynamic pages, where element-location failures decreased by 62.4%. The average maintenance effort after interface changes was reduced from 19.6 minutes to 8.3 minutes per test case. These results suggest that interaction-level planning and flakiness repair can improve the practical value of automated UI test generation.

Keywords
UI testingtest script generationweb applicationsmobile applicationssoftware testingflakiness repairmulti-agent systems
References
  1. Bao, Y., & Wang, H. (2026). Orion: A High-Throughput, Fault-Tolerant Event Routing Architecture for Hyperscale Data Streams. Fault-Tolerant Event Routing Architecture for Hyperscale Data Streams (January 01, 2026).
  2. Yaraghi, A. S., Holden, D., Kahani, N., & Briand, L. (2025). Automated test case repair using language models. IEEE Transactions on Software Engineering, 51(4), 1104-1133.
  3. Wang, S., Feng, Y., & Fang, X. (2026, May). A Large Language Model-Enabled Multi-Agent Collaboration Method for Complex Task Solving. In 2026 6th International Symposium on Computer Technology and Information Science (ISCTIS) (pp. 253-256). IEEE.
  4. Zhang, Z., Wang, J., Li, Z., Wang, Y., & Zheng, J. (2025). Anncoder: A mti-agent-based code generation and optimization model. Symmetry, 17(7), 1087.
  5. Hasan, M. M., Li, H., Fallahzadeh, E., Rajbahadur, G. K., Adams, B., & Hassan, A. E. (2026). An empirical study of testing practices in open source AI agent frameworks and agentic applications. Empirical Software Engineering, 31(5), 124.
  6. Yang, Q., & Du, Y. (2026). Predicting Rollback Risks and Controlling Gradual Rollout Traffic for Online Ranking Feature Deployments.
  7. Zhao, Z., & Welsch, R. E. (2026). Point-in-Time Financial RAG with Frozen LLMs and Market-Feedback Adaptive Retrieval. arXiv preprint arXiv:2605.31201.
  8. Yuan, Y., Huang, J., Yin, J., & Xu, T. (2026). Confidence Calibration and Semantic Error Analysis for Ambiguous Coreference Resolution in User-generated Social Media Text. Available at SSRN 7393238.
  9. Trong, M. P., Son, N. T., Giang, V. T., Viet, B. H., Hai, L. N., & Tung, D. T. (2026). A process-centric review of large language models in graphical user interface testing: architectures, lifecycle impact, and challenges. PeerJ Computer Science, 12, e3695.
  10. Zhang, Z. (2026). A Study on the Impact of Cost Presentation on Decision-making Bias and Cancellation Behavior in Online Subscriptions. Available at SSRN 7349839.
  11. Qi, C., & Qiao, X. (2026). Building and Operating a Large Scale Multi-Agent System: A Case Study from Industry. Available at SSRN 6795198.
  12. Du, Y., Liu, Q., Dong, Y., & Chen, Z. (2026). Cross-Domain Transfer and Few-Shot Recognition of Defect Images in Complex Industrial Settings.
  13. Fischer, S., & Kloihofer, W. (2026). LLM Agents for Autonomous System Testing: A Semi-structured Literature Review. In International Conference on Software Quality (pp. 147-167). Springer, Cham.
  14. Xiong, W., Guo, Y., & Zeng, Y. (2026). A Study on the Energy Efficiency of MEP Systems and Coordinated Dispatch Mechanisms for Multi-energy Systems in High-density Urban Buildings. Available at SSRN 7120518.
  15. Hong, Z., Liang, S., Xu, T., & Chen, H. (2026). A Study on Failover Verification and Recovery Objective Prediction for Cross-Region Cloud Services.
  16. Huang, J., Xu, T., Yin, J., & Yang, J. (2026). Data Quality Monitoring and Intelligent Scheduling Optimization in Financial Data Workflow Automation. Available at SSRN 7179338.
  17. Sulpovar, M., Konsynski, B. R., Kanchwala, Q., & Goodhart, G. (2026). ContextNest: Verifiable Context Governance for Autonomous AI Agent. arXiv preprint arXiv:2607.02116.
  18. Jiao, Y., Shi, T., Zhao, B., & Wang, A. (2026). Retrieval-Guided Structured Reasoning and Interpretable Representation Learning for Large-Scale Video–Language Models.
  19. Liang, S., Du, Y., Chen, W., & Liu, Z. (2026). Order Allocation and Emergency Transportation Decisions Amid Multi-Source Procurement Disruptions.
  20. Bao, Y., & Wang, H. (2026). Large-scale Metrics Migration: A Systematic Approach to Rebuilding Production Measurement Infrastructure. Available at SSRN 7185660.
  21. Muhamad, H., & Taha, M. (2026). Optimizing the Performance of Web Applications in Dynamic Network Environment: A Systematic and Comprehensive Analytical Survey. Concurrency and Computation: Practice and Experience, 38(6), e70641.
  22. Yuan, Y., Yin, J., Huang, J., & Xu, T. (2026). Weakly Supervised Anomaly Detection and Privacy Risk Scoring in Community Message Threads.
  23. Liang, S., Hsu, K. H., & Liu, Z. (2026). Rolling Scheduling and Capacity Balancing for Server Assembly Under Engineering Change Disturbances.
  24. Mabotha, E., Mabunda, N. E., & Ali, A. (2025). A dockerized approach to dynamic endpoint management for RESTful application programming interfaces in internet of things ecosystems. Sensors, 25(10), 2993.
  25. Yao, G., Zheng, J., Wang, Z., Zhang, W., Han, R., Zhao, C., ... & Liu, R. (2026, March). V-pruner: A fast and globally-informed token pruning framework for vision transformer. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 40, No. 40, pp. 34396-34404).
  26. Gu, X. (2026). Identifying Causal Effects and Analyzing Heterogeneity of User Growth Interventions on Digital Platforms: Evidence from Large-Scale Behavioral Data. Available at SSRN 6809181.
  27. Aggarwal, P., Neubig, G., & Welleck, S. (2026, April). Gym-anything: Turn any software into an agent environment. In COLM 2026 The 2nd Workshop on Lifelong Agents: Learning, Aligning, and Evolving.
  28. Zhao, J., Fan, J., & Li, L. (2026). A Study on an Explainable Causal-Enhanced LLM Agent for Predicting the Forming Quality of Automotive Component Materials.
  29. Cajas Ordóñez, S. A., Samanta, J., Suárez-Cetrulo, A. L., & Carbajo, R. S. (2025). Intelligent edge computing and machine learning: A survey of optimization and applications. Future Internet, 17(9), 417.
  30. You, S. (2026). Verifiable Audit Mechanisms in AI Compliance Automation: Scalability. Available at SSRN 6547458.
Publication details
Journal
Systems, Networks and Security Challenges
Volume
1 (2026)
Issue
1 · Forthcoming issue
Article number
snsc20260004
License
CC BY 4.0