Menu
Publications
2026
2025
2024
2023
2022
2021
2020
2019
2018
2017
2016
2015
2014
2013
2012
2011
2010
2009
2008
2007
2006
2005
2004
2003
2002
2001
Editor-in-Chief
Nikiforov
Vladimir O.
D.Sc., Prof.
Partners
doi: 10.17586/2226-1494-2026-26-3-504-516
Weakly supervised web technology identification with rule-augmented machine learning
Read the full article
Article in English
For citation:
Abstract
For citation:
Deeb R., Vorobeva A.A. Weakly supervised web technology identification with rule-augmented machine learning. Scientific and Technical Journal of Information Technologies, Mechanics and Optics, 2026, vol. 26, no. 3, pp. 504–516. doi: 10.17586/2226-1494-2026-26-3-504-516
Abstract
Identifying the technologies used by web applications is crucial for vulnerability assessment, resource inventory, and research analysis of the digital environment. Traditional tools that rely on static signature rules suffer from fundamental limitations: they require constant manual maintenance, are sensitive to code obfuscation and dynamic loading, and perform poorly when distinguishing features are subtle, distorted, or noisy. The novelty of this work lies in applying a weak supervision method based on behavioral characteristics to create a robust and scalable classifier. This work presents a hybrid approach combining weakly supervised machine learning with rule-based post-processing. Two independent signature-based systems, Wappalyzer and BuiltWith, serve as sources of weakly labeled data; a final positive label for a technology is assigned only when their detections agree. A multi-label Random Forest model is trained on this labeled data. The feature space is constructed from behavioral patterns extracted during an automated browser session, including Hypertext Transfer Protocol headers, cookies, network requests, and structural features of the Hypertext Markup Language and the Document Object Model. To correct model predictions by accounting for typical technology co-occurrences, a post-processing step based on association rules mined from the training data is applied. The experimental evaluation was conducted on a dataset of 8,594 websites for identifying 122 technologies. The Random Forest model demonstrated consistently high precision and recall for common web technologies, such as web servers, content management systems, and front-end libraries. The hybrid approach with post-processing provided a statistically significant improvement in both micro-averaged and weighted harmonic mean of precision and recall, while preserving the results interpretability. On a fixed test set (20 % of the data), the Random Forest model achieved a micro-averaged harmonic mean of precision and recall of 0.750, while the hybrid approach reached 0.763. Association rules affected 23.8 % of the predictions, moderately improving recall. For widespread technologies (Apache, Nginx, WordPress, jQuery), the harmonic mean of precision and recall exceeded 0.840, and for several components (Drupal, Google Cloud) it reached values of 0.960–1.000. The obtained results show that the approach based on weak supervision and behavioral analysis can provide a scalable method for web technology identification. It effectively complements traditional signature-based methods in cases of deliberate code concealment or dynamic execution. This method is promising for automated auditing, enhancing security tool capabilities, and scalable research of the web technology landscape. Future work may focus on improving the feature set and adapting the system for detecting novel and unique technologies.
Keywords: web technologies identification, weak supervision, machine learning, cybersecurity, behavioral analysis, Random Forest, association rules
References
References
1. Skolka P., Staicu C.-A., Pradel M. Anything to hide? studying minified and obfuscated code in the web. Proc. of the World Wide Web Conference (WWW 2019), 2019, pp. 1735–1746. doi: 10.1145/3308558.3313752
2. Sarker S., Jueckstock J., Kapravelos A. Hiding in plain site: detecting JavaScript obfuscation through concealed browser API usage. Proc. of the ACM Internet Measurement Conference, 2020, pp. 648–661. doi: 10.1145/3419394.3423616
3. Applebaum S., Gaber T., Ahmed A. Signature-based and machine-learning-based web application firewalls: a short survey. Procedia Computer Science, 2021, vol. 189, pp. 359–367. doi: 10.1016/j.procs.2021.05.105
4. Zheng R., Ma H., Wang Q., Fu J., Jiang Z. Assessing the security of campus networks: The case of seven universities. Sensors, 2021, vol. 21, no. 1, pp. 306. doi: 10.3390/s21010306
5. Zhang J., Hsieh C.-Y., Yu Y., Zhang C., Ratner A. A survey on programmatic weak supervision. arXiv, 2022. arXiv:2202.05433. doi: 10.48550/arXiv.2202.05433
6. Agrawal R., Imieliński T., Swami A. Mining association rules between sets of items in large databases. ACM SIGMOD Record, 1993, vol. 22, no. 2, pp. 207–216. doi: 10.1145/170036.170072
7. Berg A., Lamberg N. Automatic Fingerprinting of Websites. Technical report, 2020.
8. Shi Y., Yu W., Zhao Y., Jia Y. A web application fingerprint recognition method based on machine learning. Computer Modeling in Engineering & Sciences, 2024, vol. 140, no. 1, pp. 887–906. doi: 10.32604/cmes.2024.046140
9. Rizzo V., Traverso S., Mellia M. Unveiling web fingerprinting in the wild via code mining and machine learning. Proceedings on Privacy Enhancing Technologies, 2021, vol. 2021, no. 1, pp. 43–63. doi: 10.2478/popets-2021-0004
10. Qiang W., Ren K., Wu Y., Zou D., Jin H. DeepFPD: browser fingerprinting detection via deep learning with multimodal learning and attention. IEEE Transactions on Reliability, 2024, vol. 73, no. 3, pp. 1516–1528. doi: 10.1109/TR.2024.3355233
11. Bahramali A., Bozorgi A., Houmansadr A. Realistic website fingerprinting by augmenting network traces. Proc. of the ACM SIGSAC Conference on Computer and Communications Security, 2023, pp. 1035–1049. doi: 10.1145/3576915.3616639
12. Deng X., Li Q., Xu K. Robust and reliable early-stage website fingerprinting attacks via spatial-temporal distribution analysis. Proc. of the ACM SIGSAC Conference on Computer and Communications Security, 2024, pp. 1997–2011. doi: 10.1145/3658644.3670272
13. Topcuoglu C., Onarlioglu K., Jabiyev B., Kirda E. Untangle: multi-layer web server fingerprinting. Proc. of the Network and Distributed System Security Symposium, 2024, doi: 10.14722/ndss.2024.24497
14. Bird S., Mishra V., Englehardt S., Willoughby R., Zeber D., Rudametkin W., Lopatka M. Actions speak louder than words: semi-supervised learning for browser fingerprinting detection. arXiv, 2020. arXiv:2003.04463. doi: 10.48550/arXiv.2003.04463
15. Lin X., Araujo F., Taylor T., Jang J., Polakis J. Fashion faux pas: implicit stylistic fingerprints for bypassing browsers’ anti-fingerprinting defenses. Proc. of the IEEE Symposium on Security and Privacy (SP), 2023, pp. 987–1004. doi: 10.1109/SP46215.2023.10179437
16. Le Pochat V., Van Goethem T., Tajalizadehkhoob S., Korczyński M., Joosen W. Tranco: a research-oriented top sites ranking hardened against manipulation. Proc. of the 26th Annual Network and Distributed System Security Symposium, 2019, doi: 10.14722/ndss.2019.23386
17. Zhang M.-L., Zhou Z.-H. A review on multi-label learning algorithms. IEEE Transactions on Knowledge and Data Engineering, 2014, vol. 26, no. 8, pp. 1819–1837. doi: 10.1109/TKDE.2013.39
18. Altulaihan E.A., Alismail A., Frikha M. A survey on web application penetration testing. Electronics, 2023, vol. 12, no. 5, pp. 1229. doi: 10.3390/electronics12051229
19. Abusnaina A., Jang R., Khormali A., Nyang D., Mohaisen D. DFD: adversarial learning-based approach to defend against website fingerprinting. Proc. of the IEEE INFOCOM 2020 – IEEE Conference on Computer Communications, 2020, pp. 2459–2468. doi: 10.1109/INFOCOM41043.2020.9155465
20. Breiman L. Random forests. Machine Learning, 2001, vol. 45, no. 1, pp. 5–32. doi: 10.1023/A:1010933404324
21. Pedregosa F., Varoquaux G., Gramfort A., Michel V., Thirion B., Grisel O., et al. Scikit-learn: machine learning in Python. Journal of Machine Learning Research, 2011, vol. 12, pp. 2825–2830.

