SME Business Intelligence Support Using Retrieval-Augmented Generation and RFM Segmentation
Abstract
This study presents the design and evaluation of a cloud-based business intelligence support system for small and medium enterprises that integrates retrieval-grounded text generation with recency–frequency–monetary customer segmentation to enhance digital customer communication and promotional decision making. The primary objective is to assist individual small businesses in responding accurately to customer inquiries while simultaneously leveraging historical transaction data to identify actionable customer groups, all within their existing messaging workflows through a mobile keyboard interface. The proposed framework combines two complementary components. The first component automatically generates customer replies by retrieving semantically relevant information from a structured business knowledge base and using it to produce grounded, context-aware responses. The second component analyzes invoice records to segment customers into loyal, moderate, and at-risk groups, enabling sellers to tailor promotional strategies based on observed purchasing behavior. The system is implemented as a cloud service accessed by individual enterprises without requiring local infrastructure or model training. System evaluation was conducted using real small business data collected over several weeks. Experimental procedures included retrieval faithfulness analysis, response correctness evaluation with confidence intervals, customer cluster validation using silhouette analysis, end-to-end latency measurement, and structured user acceptance testing. Performance results demonstrate that the retrieval mechanism consistently provides accurate knowledge grounding, while the segmentation module effectively distinguishes high-value and churn-risk customers. The average response time remained within a range perceived as acceptable for real-time mobile conversations, and user testing confirms that the keyboard-based interface does not disrupt normal communication practices. The findings indicate that embedding retrieval-grounded generation and lightweight customer analytics directly into daily messaging tools can significantly improve the operational efficiency of small enterprises. This integrated approach reduces the burden of manual response handling while enabling data-driven promotional decision making. The framework offers a practical pathway for adopting artificial intelligence in small business environments and provides a foundation for future enhancements such as temporal behavior modeling and multilingual support.
Keywords
Full Text:
PDFReferences
S. Adomako and M. Ahsan, “Entrepreneurial passion and SMEs’ performance: Moderating effects of financial resource availability and resource flexibility,” Journal of Business Research, vol. 144, pp. 122–135, May 2022, doi: https://doi.org/10.1016/j.jbusres.2022.02.002.
Indah Ramadhani, D. Dailami, Ulva Widiya, I. Yunita, Ridha Nurhuda, and M. Mutia, “The Influence of Response Speed and Information Quality on the Effectiveness of SME Sales Through Facebook,” Golden Ratio of Mapping Idea and Literature Format, vol. 6, no. 1, pp. 117–124, Jul. 2025, doi: https://doi.org/10.52970/grmilf.v6i1.1377.
S. W. Arista, Sigit Hermawan, S. W. Arista, and Sigit Hermawan, “Improving MSME Performance Based on Digital Marketing, Intellectual Capital, Product Innovation and Competitive Advantage,” Jurnal Manajemen Bisnis, vol. 16, no. 2, pp. 526–557, Aug. 2025, doi: https://doi.org/10.18196/mb.v16i2.25192.
M. S. Mahrinasari, S. Bangsawan, and M. F. Sabri, “Local wisdom and Government’s role in strengthening the sustainable competitive advantage of creative industries,” Heliyon, vol. 10, no. 10, p. e31133, May 2024, doi: https://doi.org/10.1016/j.heliyon.2024.e31133.
F. S. Singagerda and S. Riadi, “Competitive Advantage of Family Business and the Barriers: Evidence from Indonesia SME’s,” Proceedings of the Proceedings of the 1st Workshop on Multidisciplinary and Its Applications Part 1, WMA-01 2018, 19-20 January 2018, Aceh, Indonesia, 2019, doi: https://doi.org/10.4108/eai.20-1-2018.2281916.
S. Qaiser and R. Ali, “Text Mining: Use of TF-IDF to Examine the Relevance of Words to Documents,” International Journal of Computer Applications, vol. 181, no. 1, pp. 25–29, Jul. 2018, doi: https://doi.org/10.5120/ijca2018917395.
K. Zhou, K. Ethayarajh, D. Card, and D. Jurafsky, “Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words,” arXiv.org, May 10, 2022. https://arxiv.org/abs/2205.05092 (accessed Apr. 18, 2024).
M. Ahmed, H. U. Khan, S. Iqbal, and Qutaibah Althebyan, “Automated Question Answering based on Improved TF-IDF and Cosine Similarity,” 2022 Ninth International Conference on Social Networks Analysis, Management and Security (SNAMS), pp. 1–6, Nov. 2022, doi: https://doi.org/10.1109/snams58071.2022.10062839.
Z. FU, X. SUN, Q. LIU, L. ZHOU, and J. SHU, “Achieving Efficient Cloud Search Services: Multi-Keyword Ranked Search over Encrypted Cloud Data Supporting Parallel Computing,” IEICE Transactions on Communications, vol. E98.B, no. 1, pp. 190–200, 2015, doi: https://doi.org/10.1587/transcom.e98.b.190.
Z. Xia, X. Wang, X. Sun, and Q. Wang, “A Secure and Dynamic Multi-Keyword Ranked Search Scheme over Encrypted Cloud Data,” IEEE Transactions on Parallel and Distributed Systems, vol. 27, no. 2, pp. 340–352, Feb. 2016, doi: https://doi.org/10.1109/tpds.2015.2401003.
Chen, A. Fisch, J. Weston, and A. Bordes, “Reading Wikipedia to Answer Open-Domain Questions,” Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2017, doi: https://doi.org/10.18653/v1/p17-1171.
C. Alberti, D. Andor, E. Pitler, J. Devlin, and M. Collins, “Synthetic QA Corpora Generation with Roundtrip Consistency,” arXiv (Cornell University), Jun. 2019, doi: https://doi.org/10.48550/arxiv.1906.05416.
J. Howard and S. Ruder, “Universal Language Model Fine-tuning for Text Classification,” arXiv (Cornell University), Jan. 2018, doi: https://doi.org/10.48550/arxiv.1801.06146.
M. A. Bakker et al., “Fine-tuning language models to find agreement among humans with diverse preferences,” arXiv (Cornell University), Nov. 2022, doi: https://doi.org/10.48550/arxiv.2211.15006.
C. Pornprasit and C. Tantithamthavorn, “Fine-tuning and prompt engineering for large language models-based code review automation,” Information and Software Technology, p. 107523, Jul. 2024, doi: https://doi.org/10.1016/j.infsof.2024.107523.
N. Ding et al., “Parameter-efficient fine-tuning of large-scale pre-trained language models,” Nature Machine Intelligence, vol. 5, no. 3, pp. 220–235, Mar. 2023, doi: https://doi.org/10.1038/s42256-023-00626-4.
D. M. Anisuzzaman, J. G. Malins, P. A. Friedman, and Z. I. Attia, “Fine-Tuning LLMs for Specialized Use Cases,” Mayo Clinic Proceedings: Digital Health, vol. 3, no. 1, Nov. 2024, doi: https://doi.org/10.1016/j.mcpdig.2024.11.005.
B. Mohammadi, E. Abbasnejad, Y. Qi, Q. Wu, A. Van Den Hengel, and J. Q. Shi, “Parameter-efficient action planning with large language models for vision-and-language navigation,” Pattern Recognition, vol. 172, p. 112462, Apr. 2026, doi: https://doi.org/10.1016/j.patcog.2025.112462.
L. Wang et al., “Parameter-efficient fine-tuning in large language models: a survey of methodologies,” Artificial Intelligence Review, vol. 58, no. 8, May 2025, doi: https://doi.org/10.1007/s10462-025-11236-4.
W. Fan et al., “A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models,” in KDD '24, Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Pages 6491 – 6501, doi: https://doi.org/10.1145/3637528.3671470
Patrick et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Neural Information Processing Systems, vol. 33, pp. 9459–9474, May 2020, doi: https://doi.org/10.48550/arXiv.2005.11401.
B. Kim, J. Oh, and C. Min, “Investigation on Applicability and Limitation of Cosine Similarity-Based Structural Condition Monitoring for Gageocho Offshore Structure,” Sensors, vol. 22, no. 2, p. 663, Jan. 2022, doi: https://doi.org/10.3390/s22020663.
H. Steck, C. Ekanadham, and N. Kallus, “Is Cosine-Similarity of Embeddings Really About Similarity?,” arXiv (Cornell University), Mar. 2024, doi: https://doi.org/10.1145/3589335.3651526.
L. Giray, “Prompt Engineering with ChatGPT: a Guide for Academic Writers,” Annals of Biomedical Engineering, vol. 51, pp. 2629–2633, Jun. 2023, doi: https://doi.org/10.1007/s10439-023-03272-4.
DOI: https://doi.org/10.47738/jads.v7i2.1163
Refbacks
- There are currently no refbacks.

Journal of Applied Data Sciences
| ISSN | : | 2723-6471 (Online) |
| Publisher | : | Bright Publisher |
| Website | : | http://bright-journal.org/JADS |
| : | taqwa@amikompurwokerto.ac.id (principal contact) | |
| support@bright-journal.org (technical issues) |
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0




.png)