Most Powerful Supercomputer on Earth Joins the Fight Against COVID-19

The invisible microscopic enemy “COVID-19” is now got a fearsome opponent, “Supercomputer”. Computational Biology is about get a shot in the arm in fight against SARS-COV-2 (also known as COVID-19). The titleholder of world’s most powerful supercomputer, IBM Summit” has now joined the fight against COVID-19 in addition to crowdsource distributed supercomputing network from FoldingatHome (FAH). Housed in the US Department of Energy’s Oak Ridge National Laboratory (ORNL), IBM Summit supercomputer is primarily assigned with task to solve some of the impractical or impossible tasks in fields of energy, advanced materials, human health, and artificial intelligence.

It’s precious computational time is now allocated to researchers to perform simulation at “unprecedented speed” sifting through thousands of molecules to find potentially druggable compounds that could fight against severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the coronavirus responsible for the current COVID-19 pandemic.

The Fight of Summit against COVID-19

With it’s million times computing power compared to any latest desktop or laptop, Summit allowed researchers to see results of sifting through molecules for druggable compounds in days than what would have otherwise taken months.  As of today, researchers have identified 77 small-molecule drug compounds that the ORNL said might warrant further study in the fight against the SARS-CoV-2 coronavirus, which is responsible for the COVID-19 disease outbreak [1].

The Summit helped in simulation of 8000 compounds to screen for those that will bind to the main protein “Spike” of COVID-19 rendering it unable to infest host cells. The findings are published in a paper that is yet to be “peer reviewed” and available at the preprint server “ChemRxiv”. The SARS-CVO-2 virus surfaces are covered with spikey crown-like proteins which allows the virus to bind it to human cells herein ACE2 receptor [2].  You can view the video herein to learn more.

The below diagram depicts the compound found (identified in gray colored) using Summit supercomputer simulation that binds to SARS-COV-2 spike-protein rendering it unable to infect ACE2 receptor (purple color).

Figure 1. The compound, shown in gray, was calculated to bind to the SARS-CoV-2 spike protein, shown in cyan, to prevent it from docking to the Human Angiotensin-Converting Enzyme 2, or ACE2, receptor, shown in purple. Credit: Micholas Smith/Oak Ridge National Laboratory, US Dept. of Energy. [Courtesy: ORNL]

Computational power of Summit has shortened the time for simulation in finding druggable compounds, now it is time for scientists to start further experiments. “Our results don’t mean that we have found a cure or treatment for the Wuhan coronavirus”, said Jeremy Smith. “We are very hopeful, though, that our computational findings will both inform future studies and provide a framework that experimentalists will use to further investigate these compounds. Only then will we know whether any of them exhibit the characteristics needed to mitigate this virus.”

Summit is considered the title holder of most powerful computer on earth. Housed at Oak Ridge National Laboratory (ORNL), Tennessee, the supercomputer is the size of two tennis courts and is capable of processing over 200 quadrillion calculations per second. It is used by researchers from modeling of supernova to crunching data for cancer and genetic research. Computational biology is not new but use of supercomputer of this kind for computational biology is very promising.

History of Supercomputing at IBM

The prelude to supercomputer at IBM starts in 1954 when IBM built a vacuum tube computer known as Naval Ordnance Research Calculator (NORC) for United States Navy’s Bureau of Ordnance [3]. It calculated upto 3089 digits which was a record at the time.  Since then IBM produce a series of computational systems (please refer to figure 2) with most notable first supercomputer that challenged human brain power was deep blue developed in 1997. It defeated Garry Kasparov (a chess grandmaster and world chess champion at the time) at first round of the six games match though Gary won later three games. The next notable supercomputer was Watson developed in 2011 by IBM which challenged human brain power and own first prize in jeopardy competition.

Figure 2. A Brief history of supercomputing at IBM

In 2013, IBM announced that the first commercial application of Watson for the utilization of decisions in lung cancer treatment at Memorial Sloan Kettering Cancer Center, New York City [4]. Following Watson, IBM built Sequoia supercomputer in 2012 and delivered to Lawrence Livermore National Laboratory (LLNL). It performed well against K Computer with 17.17 petaflops compared to 10.51 petaflops by K computer. Soon after Sequoia, IBM developed “Sierra” and it’s sibling “Summit” in 2018. Sierra was delivered to LLNL while Summit to ORNL. Sierra supports upto 4320 nodes for the system delivered to LLNL while Summit supports 4608 nodes. The peak performance of Sierra is 125 petaflops (PF) while performance of Summit is 200 PF. Both systems use NVIDIA GPUs and Mellanox Infiniband EDR.

Underneath the hood of Summit

The Summit supercomputer consists of 4,608 compute nodes each with two 22-core POWER9 (P9) processors and six NVIDIA Tesla V100 (Volta) GPUs. It uses 50Gb/s NVLink2.0 bus to connect each P9 CPU to three V100 GPUs and GPUs to each other as shown in figure 3.

Figure 3. Summit Compute Node Architecture (Larrea, et al., n.d.).

The CPUs are connected to 256 GB DDR4 memory and the system includes 1.6TB NVMe storage. For shared storage centralized Alpine GPFS file system (Larrea, et al., n.d.). Each compute node is stacked in a rack upto 18 x 2 RU compute nodes. The network architecture is build using flat tree topology as depicted in the figure 4.

Figure 4. Summit Supercomputer compute node and deployment architecture (Larrea, et al., n.d; Kahle & Dreps, 2019).

Both infiniband and Ethernet network connectivity provides storage and transport access respectively. The system has total of 256 rack, 10.2 PB memory and 250 Petabytes storage. 

If you are interested to learn more, please watch this video.

Reference

  1. ZDNET, 2020. IBM Summit supercomputer joins fight against COVID-19. Available online at https://www.zdnet.com/article/ibm-summit-supercomputer-joins-fight-against-covid-19/
  2. Chowdhury, D., 2020. LEND POWER OF YOUR COMPUTER TO FIGHT COVID-19. Available online at http://www.dhimanchowdhury.com/2020/03/16/lend-power-of-your-computer-to-fight-covid-19/
  3. Wikipedia, 2020. IBM Naval Ordnance Research Calculator. Available online at https://en.wikipedia.org/wiki/IBM_Naval_Ordnance_Research_Calculator .
  4. Wikipedia, 2020. Watson (computer). Available online at https://en.wikipedia.org/wiki/Watson_(computer) .
  5. Larrea, et al., n.d. Larrea, V.G.V., Joubert, W., Brim, J. M., Budiardja, D. R., Maxwell, D., Ezell, M., Zimmer, C., Boehm, S., Elwasif, W., Oral, S., Fuson, C., Pelfrey, D., Hernandez, O., Leverman, D., Hanley, J., Berrill, M. & Tharrington, A., n.d. Scaling the Summit: Deploying the World’s Fastest Supercomputer? Available online at https://www.osti.gov/servlets/purl/1561654 .
  6. Kahle, A. J., Moreno, J. & Dreps, D., 2019. Summit & Sierra: Designing AI/HPC Supercomputers. IEEE International Solid-State Circuits Conference: ISSCC 2019 / SESSION 2 / PROCESSORS / 2.1.

Lend power of your computer to fight COVID-19

Fight COVID-19 with your computer

 

Around the world countries are taking drastic measures to protect their citizens against a rare form of flu pandemic known as COVID-19 or (Corona Virus Disease 2019) [1]. While governments, researchers and health professionals working tirelessly to contain the pandemic and find cure against this rapidly spreading disease, you too can help by allocating your computer resources to towards advances in research about this rapidly spreading disease from the comfort of your home.

How it works?

Simply start “Folding” by downloading an application from Folding@home (FAH or F@h) at https://foldingathome.org/start-folding/. It will allow you to lend the power of your computer (GPU and CPU) toward advancing the research of coronavirus. Majority of today’s laptops and desktops have built-in GPU and multi-core CPU: just let the application tap into your unused clock cycles. Just connect to internet, download the application, install the FAH application and stay connected (figure 1).

Figure 1. Download and install FAH application.

Folding@Home or FAH has been around since October, 2000. It was developed by the Pande Laboratory at Stanford University, under the direction of Prof. Vijay Pande, who lead the project until 2019. Since 2019, Folding@home has been led by Dr. Greg Bowman, a former student of Dr. Pande [2]. The FAH is a distributed computing project [3] for research that simulates protein folding [4], computational drug design, and other types of molecular dynamics [5]. Till date, FAH helped Pande Lab to produce 223 scientific research papers [6]. While FAH is well known distributed computing project, there are other initiatives that uses idle computer power to carry out various research in the field of astronomy, chemistry, biology, climatology, mathematics, and physics, e.g. Berkeley Open Infrastructure for Network Computing (BOINC) [7]. Unlike FAH that focuses purpose-built for protein folding, BOINC supports 30 other science projects such as Einstein@HomeIBM World Community Grid, and SETI@home. Now that you know, you can lend your computing power to do some great goods for some of the world’s pressing problems either through FAH for COVID-19 research or through BIONC supported projects.

The concept of Distributed Computing that powers FAH and other projects to help solve world’s pressing problems

The distributed computing is a concept of using multiple computers to solve a common problem by computation distributed among connected computers unlike parallel computing systems that uses common memory pool.

Figure 2. Distributed Computing

Figure 2. Distributed Computing

There are different messaging mechanisms for distributed computing which may includes http, RPC and message queues. Implementation architecture also varies and falls into one of the following architectures:

· Client-server: It is the simplest form of architecture in which client request and server executing or some way of fulfilling the request. FAH implementation is a good example of client server architecture that uses RPC connectors for communications. A client server architecture could be two tier or three-tier depending upon the implementation.

· N-tier: This architecture also known as multi-tier mainly comprises of applications servers for which processing is performed through different computing systems as depicted in figure 2 (above).

· Peer-to-Peer (P2P): The P2P is a form of client-server distributed system is which every node can be either server or client depending upon the task. Filesharing and blockchain implementations are good examples of P2P deployments. SETI@home project uses P2P communication and computation.

· Cluster Computing: It is a “High Performance Distributed Computing” (HPDC) model in which distributed computing techniques are applied to the solution of computationally intensive applications across networks of computers. Cluster computing can be either distributed or parallel in implementation.

What FAH is doing for the research of COVID-19?

Simply put FAH is assisting researchers through simulating of druggable design of protein target to combat COVID-19. As of today, FAH has released initial wave of projects simulating potentially druggable protein targets from SARS-CoV-2 (the virus that causes COVID-19) and the related SARS-CoV virus [8]. These projects help scientist better understand how coronaviruses interact with human ACE2 receptor (also known as angiotensin converting enzyme 2). According to studies of patients with severe acute respiratory syndrome (SARS) demonstrated that the respiratory tract is a major site of SARS-coronavirus (CoV) infection and disease morbidity for which ACE2 is the viral entry point to human cell [8; 9]. Similar observations are also made for the virus infections of COVID-19. Hence, work at FAH is of great importance because FAH initiative helps researchers better understand how COVID19 interact with ACE2 and how researchers might be able to interfere with them through the design of new therapeutic antibodies or small molecules that might disrupt their interaction.

With collaboration and partnership of various labs, researchers and some of the new structural biology and biochemical data available at bioRxiv and chemRxiv, Folding@home hopes to help researchers better understand COVID-19 and how it interact with ACE2 receptor in a process to fight the virus. For further details, please read https://foldingathome.org/2020/03/10/covid19-update/ .

Back to COVID-19: computational biology in a nutshell

As discussed, COVID-19 has much similarities with SARS coronavirus from 2003. The infection resulting from both viruses occurs at lung when a virus protein (spike protein) binds with a lung cell receptor protein (ACE2) [8;9]. An antibody can prevent the spread of the infection by blocking the spike protein from binding with the receptor, in other words, blocking the COVID-19 virus to bind ACE2. To develop such antibody, researchers need to study the structure of COVID-19 spike protein, various shapes it takes and how it binds to ACE2 receptor. This requires lots of protein folding simulations unique to COVID-19 and that where FAH comes in. Protein folding is constantly ongoing and a sorts of biological opera with a huge cast of performers, an intricate plot, and dramatic denouements when things go awry [10]. Understanding these intricacies help scientist create druggable protein design to fight disease among other things. Accurate simulation of protein folding is thus holy grail of computational biology. In a human body, protein takes milliseconds to seconds to fold completely. For a computer to simulate this protein folding with all changes at atomic level, it has to perform it has to perform trillions of quadrillions of steps. If we were to simulate this using single high end personal computer, it will take centuries to render. Please view the following video on millisecond protein folding to better understand how things work.

To render similar tasks of a millisecond-level simulation of what you are viewing in this video above, may take took months of CPU time for a typical “Supercomputer”. That’s where distributed computing model such as Folding@home (FAH) is useful. The FAH computing model behaves as “Supercomputer” by using power of computers from volunteers around the world and performance is also great with more than 100 Peta FLOPS (floating point operation per second) in comparison to IBM supercomputer with 150 PetaFLOPS: FLOPS is a unit to measure the performance of “Supercomputer”. Additionally, FAH deployment is much cheaper than using IBM’s OLCF-4 type supercomputer. More importantly, according to the researchers in the Pande Lab, “Protein folding dynamics is statistical in nature, so a single long simulation from a supercomputer would not be sufficient to fully understand the folding process” [11].

Reference:

1. CDC, 2020. Key facts: Know the facts about coronavirus disease 2019 (COVID-19) and help stop the spread of rumors. National Center for Immunization and Respiratory Diseases (NCIRD), Division of Viral Diseases. Available online at https://www.cdc.gov/coronavirus/2019-ncov/symptoms-testing/share-facts.html?CDC_AA_refVal=https%3A%2F%2Fwww.cdc.gov%2Fcoronavirus%2F2019-ncov%2Fabout%2Fshare-facts.html

2. Wikipedia, 2020. Folding@home. Wikipedia. Available at https://en.wikipedia.org/wiki/Folding@home.

3. ScienceDirect, 2020. Distributed Computing. Science Direct. Available at https://www.sciencedirect.com/topics/computer-science/distributed-computing

4. Wikipedia, 2020. Protein Folding. Wikipedia. https://en.wikipedia.org/wiki/Protein_folding .

5. Meller, J. 2001. Molecular Dynamics. Encyclopedia of Life Sciences: Nature Publishing Group. Available https://dasher.wustl.edu/chem430/readings/md-intro-1.pdf .

6. FAH, 2020. Papers and Results. Folding@Home. Available online at https://foldingathome.org/papers-results/

7. BOINC, 2020. Compute for Science. BOINC. Available at https://boinc.berkeley.edu/ .

8. FAH, 2020. FOLDING@HOME UPDATE ON SARS-COV-2 (10 MAR 2020). Folding@Home. Available https://foldingathome.org/2020/03/10/covid19-update/ .

9. Jia et al, 2005. Jia, P.H., Look, C.D., Shi, L., Hickey, M., Pewe, L., Netland, J., Farzan, M., Wohlford-Lenane, C., Perlman, S. & McCray, B. P., 2005. ACE2 Receptor Expression and Severe Acute Respiratory Syndrome Coronavirus Infection Depend on Differentiation of Human Airway Epithelia. Journal of Virology, American Society for Microbiology (ASM).

10. Everts, S., 2017. Protein folding: Much more intricate than we thought. C & EN. Available at https://cen.acs.org/articles/95/i31/Protein-folding-Much-intricate-thought.html .

11. Mathi, S. 2020. You Can Help Fight Coronavirus by Giving Scientists Access to Your Computer. Available at https://onezero.medium.com/you-can-help-fight-coronavirus-by-giving-scientists-access-to-your-computer-16c39c2e7164 .

Disaggregation brightens the prospect for Edge

Disaggregation brightens the prospect for Edge

Today, cloud and virtualization are increasingly pushing the need to abolish hierarchical construct of network to a simpler and dynamic flat service-oriented infrastructure. And, it is for good reason: create non-blocking architecture, reduce cost and complexity and improve latency and manageability of the network. As network flattened, tiered notion of aggregation blends away to form two distinct elements of network: Edge and Core. While Edge extend service-oriented infrastructure to network entry point, network core multiplexes zeta bytes of data transport making it easier to traverse multiple operators in a unique mesh architecture.

Figure 1. An example of Flat network topology that integrate distributed Data centers to a service-oriented edge network.
Figure 1. An example of Flat network topology that integrate distributed Data centers to a service-oriented edge network.

The above figure depicts example of service-oriented edge network which integrates mobile, fixed wireless access (FWA) and fixed line services to distributed datacenters (DC) eliminating the need for centralized data processing. In such architecture data processing is localized and only certain data elements are retrieved from regional or central DCs as required and on demand without bogging down the pipe. The infrastructure topology presented here depicts how fixed line and wireless access services are combined and extended over long distance using open optical packet transponder (OOPT).

What is Edge and Why it is creating a paradigm shift?

Let use first understand two distinct parts of Edge: a service-oriented network infrastructure and computing or processing element for which later is known as Edge computing. Sometimes, pandit refers both elements of Edge as Edge computing but it innately flawed. Edge computing is enabled through Edge network infrastructure. When I am referring to Edge, it is about service-oriented network infrastructure that brings enhance experience to users at the network entry point. For this abstraction, Edge Computing is a complimentary subset than the entirety.

A service-oriented infrastructure cannot afford processing delay, increase latency and vendor locked network. More importantly, service-oriented infrastructure requires low cost per bit to innately improve service for enhance user experience. Cloud providers has done well reducing cost per processing unit to deliver enhance service. But for operators, network infrastructure remains rigid and costly making it difficult to provide enhanced services to improve their bottom-line. This is about to change, thanks the notion of flat topology and disaggregation.

That being said, Edge is not where you think it is. A simple contemplation of edge is to understand it from the perspective of network entry point to the provider network. Whether network entry point is at wireless endpoint or fixed line network. For the former, a network endpoint is fronthaul whereas later can be WAN Edge.

Our desire to stay connected, make sense out of our physical world and enhance our experience of things around us has created interesting business challenges for service providers to create values, gain business insights and provide enhance services from connected things and data that are generated from those devices. This value creation allows service provider to offer enhance human experience through various services for plethora of applications e.g. smart home, smart city, smart health, Smart retail, Smart farming and Smart Industry etc.

These process of value creation and service endowment at Edge are presented in the figure below as a layered concept. It is expected that 1 Trillion sensors and 75 billion devices will be connected to provider network by 2025: Some through radio frequency and others through intermediate gateways. To get this into perspective, just imagine that for every autonomous vehicle may have 200 sensors for vehicle. As we build smart application around the concept of smart everything from home to cities, estimate of connected device provided here above is no exaggeration.

Through this enormous interconnect we are bringing technology to solve world’s pressing problems from traffic engineering for congested cities to managing home appliance and securing our Homes. Possibilities are endless!

Figure 2. Layered concept of Edge.
Figure 2. Layered concept of Edge.

Therefore, unless network Edge is created in a manner to serve our desire for a connected society and enhance human experience, attempts to provide such services will be ineffective. There lies the benefit of Network Edge, an essential element of our connected world. This is to say, even to gain slightest benefit of technology for the betterment of mankind is nearly impossible without network Edge. The fast pace of technological advents are pushing the connectivity and value creation further. We are in for a new era, an era of connect world where technology is bringing about the progress of mankind with a rude surprise, Singularity!

Get the picture!

That’s why Edge is important and it is about to get a shot in the arm. Thanks to disaggregation that allows service provider to eliminate constraints of rigid vendor locked network to simpler, flat and service oriented infrastructure.

How Disaggregation helps?

Network is considered laggard when comes to interoperability, bandwidth on demand, improve latency, cost and auto provisioning or for better programmability. While computing and even storage made a leap of how convergence and virtualization can benefit, network till date is questionably rigid and slow. It was not until few years back when hyperscale data center decided that it is time for a second look that rigidity of network started to diminish giving birth to a relatively programmable, cost effective multi-vendor network. Industry pondered over it to bring about the change 1990s notion of network giving birth to a new concept of networking, “Open Networks”. Today, this notion of network is gaining momentum, industry is bringing about collective innovation at fore to enable the possibility of better service-oriented network that innately capable of bandwidth on-demand and flexible to accommodate what future holds with enormous cost savings. There is great interest among operators to explore and build cost effective, flat and service-oriented on-demand infrastructure and work is underway to upgrade their network. Service provider networks are still laggard compared to cloud providers or data center operators. Multitude of issues has plagued service provider network that stems from lack of technological solution and regulatory requirements to integrate and virtualize network functions creating on-demand and programmable network with low cost per bit.

It is about to change now, disaggregation that decouples hardware and software now making it possible for service provider to benefits from industry innovation, experiment new service technologies and create values from data that traverse through their network.

Figure 3. A diagrammatical representation of disaggregation and typical placement of disaggregated system in operator network.
Figure 3. A diagrammatical representation of disaggregation and typical placement of disaggregated system in operator network.

Disaggregation is opening up both wireless and fixed network entry point while reducing network latency and pushing Ethernet and variation thereof all the way to network entry point. Imagine, a pipe of 1GbE that providing 4G/LTE service can now have the possibility to become enormously fattened upto 100GbE way up the tower where RF meets digital bitstream. You heard of 5G that innately create great possibilities but do you know as such is only possible through the improve latency and enormous fat pipe that is the result of collective industry innovation. When I say “collective” it is not about individual organization building their own technologies rather a notion of industry coming together to create technological advent both in hardware and software.

This is where disaggregation finds its importance, a notion that allows hardware and software vendors to bring about best they could offer while working within the context of “togetherness” bringing the best breed of technologies to solve industry’s pressing problem.

Figure 4. Industry momentum towards disaggregation in Edge.
Figure 4. Industry momentum towards disaggregation in Edge.

This “coming together” to solve complex network issues is gaining momentum as evident from the figure above. Industry forums and organizations are working together to create open standards of multi-vendor network. Much of the industry’s pressing problem are getting solved as a result, e.g. disaggregation for cell site gateway initiative created DCSG (Disaggregated Cell Site Gateway) for mobile backhaul aggregation, Open Optical Packet Transponder (OOPT) for long distance optical transport and Open RAN (Radio Access Network) for fronthaul. Consecutively, disaggregation has penetrated in service provider Edge (PE) and Core in addition to Data centers.

What’s next?

Possibilities are endless, with AI/ML coming to picture a network can be self-healed, state aware and smart in future. Some refers this as Cognitive networks others call it intelligent networks. Irrespective of the terms, future of networks offers great promises. Disaggregation is just the beginning and it has created potentials for programmable state aware networks to be realized. It is now upto the provider and enterprise to maximize benefit.

768K Day – Internet Doomsday? Is it real?

There is an ominous rumbling in the internet about 768K day, some even termed it internet doomsday others called it “Y2K” of internet. The fear is justified given the experience of wide spread internet outage during 512K day when internet BGP table size exceeded 512,000 routes. The 512K day caused havoc and many routers simply exhausted of TCAM (Ternary content-addressable memory) size and were unable to process certain routes leaving parts of internet unreachable. The same issue seems possible this year again when internet routes exceeds 768K routes. Some predicts August 12 is 768K day following 512k Day which happened in August 12, 2014[1].

Some Internet Outages are predicted

While majority of teir1 ISPs were caught off guard during 512K day, this should not be the case this time around. There are mechanisms within BGP route configuration to protect routers from exhausting TCAM and presumably ISPs upgraded their legacy routers with patches. However, this is temporary fix and may come with string attached on reachability etc. A more permanent fix requires routers to accept upto full internet routes in its forwarding table. This requires a special memory known as “TCAM”. Older routers are built with limited TCAM size and hence, unable to provide faster response and may exhaust its resources and eventually fail or unable to process certain prefixes.

A typical outage could be something similar to the following diagram.

Figure 1. [Courtesy ThousandEyes] This diagram is presented in the blog page of “Thousand Eyes” at https://blog.thousandeyes.com/what-is-768k-day/ to depict recent outage in bayarea.
A blog post in “Thousand Eyes”[2] claimed that the writer has observed packet lost on several interfaces in the Cogent (AS 174) network in San Francisco Bayarea. As a result, many peer ISPs like comcast, quest, amazon, 8×8 etc were affected. Recently, several media reported outage in Australia in which media outlet like CSO[2] and Computerworld[3] claimed the outage directly related to BGP prefixes reaching 768K. There seems to be an increase phenomena of internet outage reported in twitter[5] message. Proper analyses could ascertain how many of reported outages are due to BGP prefix issues. Nonetheless, BGP route sizes are increasing and this will definitely cause network reachability issues in routers that lacks bigger forwarding table.

According to CIDR report (a website that keeps track of global BGP routes), the BGP route size already exceeded 768K and current table size shown as 783K as of June 11th, 2019[6]. However, this report is not official and may include duplicates.
Irrespective of the numbers presented in CIDR report, I can ascertain that majority of the customers I talked to, are looking to replace or add edge routers with 750k+ table size for IPV4 and around 65K for IPv6. Henceforth, it should go without saying that one should be cognizant of the issue and take precaution before 768K day arrives.

Under the hood

Internet routers generally process route request in two tables in conjunction with routing protocols: RIB and FIB. While RIB is part of control plane and generally processed by NOS, much of the table look and processing are done at FIB level which part of Routing hardware or reside within the pipeline of Merchant Silicon or ASIC.

Figure 2. Routing functions including RIB and FIB processing.

If the lookup engine (TCAM) within ASIC pipeline lacks capability of processing certain number of tables for IPv4 packets, RIB may flood the table causing overflow problem. With patches, routers may able to control processing and lookups at ASIC pipeline. However, such patches are temporary fix and protects routers from failing. A more permanent fix is to somehow connect ASIC packet processing pipeline to external TCAMs using high speed bus. Older ASICs lacks such capabilities resulting routers more software dependable and may be limited in table size capabilities.

Solution: Can whitebox Switch help?

Disaggregation or whitebox is the best solutions for this problem. Buy the choice of hardware herein router/switch from your preferred vendor and select software or Network Operating System (NOS). The benefit of whitebox or disaggregation for that matter allows you to select best of the breed merchant silicon and buy those in your terms with a price point you can afford. Result you get the best of both world: hardware and software. For the bigger IPv4 table size, Broadcom® Qumran-MX™ silicon with BCM52311™ Knowledge-Based Processor (KBP) [TCAM] provides you optimal choice for upto 1 million IPv4 routes. There are cases where upto 1.2 million routes are possible in such system.

Figure 3. Edge Router Whitebox based on Broadcom® Qumran-MX.

A number of hardware vendors are currently offering Qumran-MX based platform with industry proven NOS from companies such as IP Infusion.  As depicted in the figure above, Merchant Silicon herein Broadcom® Qumran-MX™ is connected through an internal bus (known as ELK bus) to external TCAM which provides further capabilities for lookups.

However, it is also important to select appropriate software vendor that has optimize such boxes and provides optimal route capabilities of more than 768K to facilitate your upgrade or help in your preparation for 768K day.

IPInfusion’s OcNOS™ is tested with a number of Hardware vendors providing you a wide slections, please ensure you select appropriate Qumran-MX based hardware with External TCAM to achieve upto 1 Million route. If you are interested about OcNOS and how it can solve your 768K day, you may visit their website at https://www.ipinfusion.com .

However, please make sure you ask each vendor to provide you with test report or atleast enough data to make educated decision.

 

Reference

[1] Some internet outages predicted for the coming month as ‘768k Day’ approaches. Available at https://www.zdnet.com/article/some-internet-outages-predicted-for-the-coming-month-as-768k-day-approaches/.

[2] Australian Internet Users Face Looming ‘768k Day’. Available at https://www.cso.com.au/mediareleases/34669/australian-internet-users-face-looming-768k-day/

[3] Australian Internet Users Face Looming ‘768k Day’. https://www.computerworld.com.au/mediareleases/34669/australian-internet-users-face-looming-768k-day/.

[4] Thousand Eyes. What is 768K Day, and Will It Cause Internet Outages? Available at https://blog.thousandeyes.com/what-is-768k-day/.

[5] Internet outage tag at twitter. Available at https://twitter.com/search?q=internet%20outage&src=tyah

[6] CIDR, 2019. CIDR report for June 11, 2019. Available at https://www.cidr-report.org/as2.0/

Breaking the barrier to Service

Today, most businesses don’t need any convincing argument about how digital transformation can enhance their agility. The pressing question is how to onboard more and more consumers and how to provide improve service through differentiated offerings.

Given the pivotal role of networks in this era of cloud-powered enterprise, low-cost-per bit is central to creating digital transformation. This imperative is equally applicable from edge of the network to the core, as well the very data center that fuels  services.

The need for low-cost-per-bit at the Edge of Network

As applications are quizzing the network, demand for bandwidth is growing. Multi-media transport is a reality today and with that reality is the need  for faster response time. Service providers are now baffled with a two-prone challenge: How to improve differentiated service with faster response time and reduce the cost to service. Consumers are demanding better service at relatively cheaper price points for multi-media payload that fuels their need to stay connected, pushing the Petabytes of data to Internet that traverse through various service provider networks. The age-old infrastructure of service provider is simply inadequate to accommodate this demand for more bandwidth. Couple this with the traffic generated by billions of connected devices. Get the picture?

Welcome to the dawn of Zetabyte era! “If each petabyte in a zettabyte were a centimeter it would mean one can reach 12 times higher than world’s tallest tower Burj Khalifa”.[1]

Accommodating this pressing need for bandwidth means opening up the clogged pipe at the wireless and wired edge and web-scale network performance at the Edge of Network. There is no better technology than “Ethernet” and “Fiber” to open up clogged pipe at wired and Wireless Edge.

Unchoking Wireless Edge

You heard it right, wireless network across the globe is getting a shot in the arm as service providers are upgrading the first choking block where your wireless traffic runs down the tower to meet wired pipe. Perhaps you too have experienced this change to some extent while browsing Internet or watching YouTube video, there is no screeching halt or buffering of app in your smartphone. While this experience may be limited to users of certain service providers,  work is underway for all tiers of providers to unchoke their wireless edge. It is a dire need of their business.

When your data rides over RF frequency from your cell phone to reach the antenna of your mobile provider, it undergoes digitization at an “analog to digital” conversion device known as RRH (Radio Remote Head). From there, the bitstreams run down the pipe to meet its first transport gateway of packetization known as Baseband Unit or short BBU. Up until this point, the network is known as fronthaul. Please read more in my earlier article to know more about frounthaul at http://www.dhimanchowdhury.com/2018/01/13/hello-world/ .

In early days, constraints of interconnects choked this entry point of your data. However, new technologies made it possible to reshape the architecture of interconnects and moves the bulky pieces of equipment away from the tower base. Fiber is now rising up the tower upto RRH and allowing bigger pipe up to 20GbE to antenna base through a means of interconnect known as CPRI. Compare this to early days of coaxial interconnect that only allowed short distance interconnect and less than 100MBps of transfer rate. In my article about wireless and fiber convergence, I made a case for overhauling RRH with addition of Ethernet. If CPRI is limited to 20GbE bandwidth, having Ethernet at the very base of antenna the pipe could be enormous: 100GbE to 400GbE. IEEE already created an open and free standard to create ethernet transport at fronthaul known as ROE (Radio over Ethernet). While RoE promises the potential for fatter pipe at cheaper cost, it is still in research phase. What possible though is the bandwidth upgrade in the connectivity between BBU and operator’s packet network known as “mobile backhaul”.

Figure 1. Typical fronthaul and backhaul Network architecture

Many operators started to upgrade the pipe between BBU and backhaul aggregation using a device known as cell site router or sometimes referred to as cell site gateway. This creates a fatter pipe from 1GbE to 10GbE for each fronthaul endpoint. Prior to such upgrade many deployments had 10/100Mbps at this level connectivity. Now the 4G fronthaul service can use 1GbE pipe while 5G 10GbE. This is just an example of upgrade underway. The important conduit to this upgrade is whitebox open networking platform. As such cell site router not only endows the benefits of an unlocked system, it innately reduced the cost per bit and improve services.

Unleashing the power of open networking at Wireless Edge

The Cell Site Router (please refer to figure 1) is not a new concept, there are many deployments of it using OEM switches. However, the cost of OEM switches is extremely high. It also poses challenge that inherent of vendor locked system making it difficult for operator to deploy in massive scale. This constraint can be completely eliminated by deploying “Whitebox Cell Site Router” that utilizes state of the art merchant silicon while offering the benefit of low cost per bit and the choice of hardware and software. The TIP (Telecom Infra Project) initiative for open networking infrastructure product to serve telecom service provider has undertaken activities around the same concept known as DCSG (Disaggregated Cell Site Router)[2]. The term “whitebox CSR (Cell Site Router)” and DCSG are synonymous. Central to this concept is separation of hardware and software as in whitebox but applied to Cell Site Router for specific deployments in mobile backhaul. There can be variations of DCSG devices ranging from 12 ports of 1/10GbE to 48 ports of 10GbE with 2 to 6 ports of 100GbE.

Figure 2. Whitebox Cell Site Router or DCSG (Disaggregated Cell Site Router) Architecture.

Additionally, the DCSG or Whitebox CSR provides SyncE and PTP (1588) clocking capabilities. It will help synchronize between GPS clock and packetized clocking information retrieved from packet network. Clocking is essential part of mobile backhaul transport. There are varied deployments using DCSG or whitebox CSR ranging from Ring interconnect to CLOS fabric interconnects. For example, the figure below shows typical ring interconnect allowing various RAN (Radio Access Network) configurations in MPLS network. Such configuration allows wider area coverage for various RANs including 5G and migration path from 4G to 5G for operators.

Figure 3. Typical Ring topology of interconnect for mobile backhaul.

For migration path from 4G to 5G, network operator benefits from utilizing same DCSG based network for both 4G and 5G. As shown in figure 3, 5G base station Distributed Unit (DU) [3][4] can be connected to same DCSG where 4G base station is connected. 5G deployment requires a spit of base station in two parts distributed unit (DU) and Central Unit (CU) [3][4] for which DU may reside on or near the tower and CU is placed in 5G core network. The CU is then connected to servers to provide various network functions through NFV (Network Function Virtualization). Please note DU can provide upto 100GbE link and RoE capabilities making it easier to integrate into your packet network. Within the DCSG based MPLS network primary and secondary path can be created easily making fail over faster and reliable. Other variation of this network could be DWDM based long distance interconnect between MPLS mobile backhaul (BMH) Ring and MBH aggregation (please refer figure 3).

Passive Optical Network for Wireless Edge

With the advent of whitebox OLT (Optical Line Terminator), Passive Optical Network (PON) can now be extended to fronthaul. The benefit of such deployment is that same fiber can carry mixed set of traffic allowing fronthaul connection and residential/business fixed line services over the same fiber. As such overall TCO (total cost of ownership) will be drastically reduced.

Figure 4. Mixed Traffic PON for MBH and fixed line services.

In figure 4, a typical example of mixed traffic PON deployment is shown. There can be two types of OLT deployments using whitebox solutions: whitebox OLT device and whitebox Ethernet switch with OLT transceivers.

For the first, OLT functions are realized through virtualized solution known as VOLTHA. In this scenario, mixed traffic from MBH, residential and enterprises are carried over single fiber depending upon split ratio. Typical split ratio is 1:32 meaning 32 ONUs can be connected to single incoming fiber. For the second solution, OLT transceivers can be used with standard whitebox ethernet switch. These OLT transceivers are available in SFP+ form factor making it easier to plug it in a 10GbE SFP+ port socket of the whitebox switch. OLT functions can be virtualized with VOLTHA or it can be implemented in the Ethernet switch itself. IP Infusion, a leader of independent NOS (Network Operating System) for whitebox platform, implements OLT functions through a container within its OcNOS™ software. Third-party OMCI stack that provides OLT functionalities can be integrated within OcNOS container for which Netconf API provides EMS (Element Management System) based configuration and management capabilities including provision of ONUs.

Solutions such as these reduce overall TCO for operators and offers the choices and benefits of open networking. Low-cost-per-bit is essential today for operators to improve service and maximize ROI.

 

Reference:

  1. Cisco, 2016. The Zettabyte Era Officially Begins (How Much is That?). Available at https://blogs.cisco.com/sp/the-zettabyte-era-officially-begins-how-much-is-that
  2. Lightwave, 2018. Telecom Infra Project targets disaggregated cell site gateways. Available at https://www.lightwaveonline.com/articles/2018/10/telecom-infra-project-targets-disaggregated-cell-site-gateways.html .
  3. Microsemi, 2019. 5G LTE Base stations. Available at https://www.microsemi.com/applications/5g-mobile-infrastructure/5g-lte-base-stations .
  4. Techplayon, 2017. 5G NR gNB Logical Architecture and Its Functional Split Options. Available at http://www.techplayon.com/5g-nr-gnb-logical-architecture-functional-split-options/