欢迎访问中国科学院大学学报,今天是
电子信息与计算机科学

无线密集网络中的低损耗多臂老虎机算法

  • 赵耀 ,
  • 罗喜良
展开
  • 1. 上海科技大学信息科学与技术学院, 上海 201210;
    2. 中国科学院上海微系统与信息技术研究所, 上海 200050;
    3 中国科学院大学, 北京 100049

收稿日期: 2020-01-14

  修回日期: 2020-04-20

  网络出版日期: 2020-04-20

基金资助

国家自然科学基金(61971286)资助

A low cost multi-armed bandit algorithm for dense wireless network

  • ZHAO Yao ,
  • LUO Xiliang
Expand
  • 1 School of Information Science and Technology, ShanghaiTech University, Shanghai 201210, China;
    2 Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences, Shanghai 200050, China;
    3 University of Chinese Academy of Sciences, Beijing 100049, China

Received date: 2020-01-14

  Revised date: 2020-04-20

  Online published: 2020-04-20

摘要

近年来人们对移动无线服务的需求与日俱增,为应对这一挑战,超密集无线网络被认为是下一代无线通信网络的基础设施架构和重要组成部分,基站的密集布置可以减少每个小区的服务用户数量,从而可为网络用户提供高速且低延迟的无线服务。但同时带来的不可避免的问题是用户在选择接入时会触发频繁的网络切换以确保可以接入到服务最佳的网络。用户接入问题往往被建模成在线学习模型。本文旨在寻找一个高效的在线用户接入方案以应对频繁网络切换造成的额外性能损失。通过对多臂老虎机模型的分析,提出基于操作杆淘汰机制的改进算法,并通过严格理论分析及数值仿真实验两个角度论证该算法的有效性。

本文引用格式

赵耀 , 罗喜良 . 无线密集网络中的低损耗多臂老虎机算法[J]. 中国科学院大学学报, 2022 , 39(3) : 403 -409 . DOI: 10.7523/j.ucas.2020.0011

Abstract

In recent years, people's demand for mobile wireless services has been increasing. In order to meet this challenge, ultra-dense wireless networks are considered to be the infrastructure and important components of the next-generation wireless communication network. Massive deployment of small base stations can reduce the number of network users in each cell, which can in turn provide the users with high-speed and low-latency wireless service. However, the inevitable problem brought with it at the same time is that users will cause frequent network handover when choosing access to ensure that they can access the network with the best service provider. User association problem is often modeled as the online learning model. This paper aims to find an efficient online user association scheme to deal with the additional network performance loss caused by frequent handover. Based on the analysis of the multi-armed bandit (MAB) model, this paper proposes an improved algorithm based on the arm elimination strategy, and demonstrates the effectiveness of the algorithm through rigorous theoretical analysis and numerical simulation experiments.

参考文献

[1] Boccardi F, Heath R W, Lozano A, et al. Five disruptive technology directions for 5G[J]. IEEE Communications Magazine, 1996, 52(2): 74-80.DOI:10.1109/MCOM.2014.6736746.
[2] Lee W, Cho D H. Enhanced group handover scheme in multiaccess networks[J]. IEEE Transactions on Vehicular Technology, 2011, 60(5): 2389-2395.DOI:10.1109/TVT.2011.2140386.
[3] Fischione C, Athanasiou G, Santucci F. Dynamic optimization of generalized least squares handover algorithms[J]. IEEE Transactions on Wireless Communications, 2014,13(3): 1235-1249.DOI:10.1109/TWC.2014.013014.121720.
[4] Guidolin F, Pappalardo I, Zanella A, et al. Context-aware handover policies in HetNets[J]. IEEE Transactions on Wireless Communications, 2016, 15(3): 1895-1906.DOI:10.1109/TWC.2015.2496958.
[5] Ye Q Y, Rong B Y, Chen Y D, et al. User association for load balancing in heterogeneous cellular networks[J]. IEEE Transactions on Wireless Communications, 2013, 12(6): 2706-2716.DOI:10.1109/TWC.2013.040413.120676.
[6] Videv S, Haas H. Energy-efficient scheduling and bandwidth-energy efficiency trade-off with low load [C]//2011 IEEE International Conference on Communications(ICC). June 5-9, 2011,Kyoto, Japan. IEEE, 2011: 1-5.DOI:10.1109/icc.2011.5962571.
[7] 孟庆民,赵媛媛,岳文静,等. 动态超密集网络中的Markov预测切换[J].通信学报, 2018, 39(10): 166-174.DOI:10.11959/j.issn.1000-436x.2018225.
[8] Wang Z, Li L H, Xu Y, et al. Handover control in wireless systems via asynchronous multiuser deep reinforcement learning[J]. IEEE Internet of Things Journal, 2018, 5(6): 4296-4307.DOI:10.1109/jiot.2018.2848295.
[9] Shen C, van der Schaar M. A learning approach to frequent handover mitigations in 3GPP mobility protocols [C]//2017 IEEE Wireless Communications and Networking Conference. March 19-22, 2017, San Francisco, CA, USA. IEEE, 2017: 1-6.DOI:10.1109/WCNC.2017.7925950.
[10] Zhou Y M, Shen C, van der Schaar M. A non-stationary online learning approach to mobility management[J]. IEEE Transactions on Wireless Communications, 2019, 18(2): 1434-1446.DOI:10.1109/TWC.2019.2893168.
[11] Bubeck S. Regret analysis of stochastic and nonstochastic multi-armed bandit problems[M]. Boston: Now Publishers Inc, 2012.DOI:10.1561/9781601986276.
[12] Auer P, Cesa-Bianchi N, Fischer P. Finite-time analysis of the multiarmed bandit problem[J]. Machine Learning, 2002,47: 235-256.DOI:10.1023/A:1013689704352.
[13] Sun Y X, Zhou S, Xu J. EMM: energy-aware mobility management for mobile edge computing in ultra dense networks[J]. IEEE Journal on Selected Areas in Communications, 2017, 35(11): 2637-2646.DOI:10.1109/JSAC.2017.2760160.
[14] Auer P, Ortner R. UCB revisited: improved regret bounds for the stochastic multi-armed bandit problem[J]. Periodica Mathematica Hungarica, 2010, 61(1/2): 55-65.DOI:10.1007/S10998-010-3055.6.
文章导航

/