768K Day – Internet Doomsday? Is it real?

There is an ominous rumbling in the internet about 768K day, some even termed it internet doomsday others called it “Y2K” of internet. The fear is justified given the experience of wide spread internet outage during 512K day when internet BGP table size exceeded 512,000 routes. The 512K day caused havoc and many routers simply exhausted of TCAM (Ternary content-addressable memory) size and were unable to process certain routes leaving parts of internet unreachable. The same issue seems possible this year again when internet routes exceeds 768K routes. Some predicts August 12 is 768K day following 512k Day which happened in August 12, 2014[1].

Some Internet Outages are predicted

While majority of teir1 ISPs were caught off guard during 512K day, this should not be the case this time around. There are mechanisms within BGP route configuration to protect routers from exhausting TCAM and presumably ISPs upgraded their legacy routers with patches. However, this is temporary fix and may come with string attached on reachability etc. A more permanent fix requires routers to accept upto full internet routes in its forwarding table. This requires a special memory known as “TCAM”. Older routers are built with limited TCAM size and hence, unable to provide faster response and may exhaust its resources and eventually fail or unable to process certain prefixes.

A typical outage could be something similar to the following diagram.

Figure 1. [Courtesy ThousandEyes] This diagram is presented in the blog page of “Thousand Eyes” at https://blog.thousandeyes.com/what-is-768k-day/ to depict recent outage in bayarea.
A blog post in “Thousand Eyes”[2] claimed that the writer has observed packet lost on several interfaces in the Cogent (AS 174) network in San Francisco Bayarea. As a result, many peer ISPs like comcast, quest, amazon, 8×8 etc were affected. Recently, several media reported outage in Australia in which media outlet like CSO[2] and Computerworld[3] claimed the outage directly related to BGP prefixes reaching 768K. There seems to be an increase phenomena of internet outage reported in twitter[5] message. Proper analyses could ascertain how many of reported outages are due to BGP prefix issues. Nonetheless, BGP route sizes are increasing and this will definitely cause network reachability issues in routers that lacks bigger forwarding table.

According to CIDR report (a website that keeps track of global BGP routes), the BGP route size already exceeded 768K and current table size shown as 783K as of June 11th, 2019[6]. However, this report is not official and may include duplicates.
Irrespective of the numbers presented in CIDR report, I can ascertain that majority of the customers I talked to, are looking to replace or add edge routers with 750k+ table size for IPV4 and around 65K for IPv6. Henceforth, it should go without saying that one should be cognizant of the issue and take precaution before 768K day arrives.

Under the hood

Internet routers generally process route request in two tables in conjunction with routing protocols: RIB and FIB. While RIB is part of control plane and generally processed by NOS, much of the table look and processing are done at FIB level which part of Routing hardware or reside within the pipeline of Merchant Silicon or ASIC.

Figure 2. Routing functions including RIB and FIB processing.

If the lookup engine (TCAM) within ASIC pipeline lacks capability of processing certain number of tables for IPv4 packets, RIB may flood the table causing overflow problem. With patches, routers may able to control processing and lookups at ASIC pipeline. However, such patches are temporary fix and protects routers from failing. A more permanent fix is to somehow connect ASIC packet processing pipeline to external TCAMs using high speed bus. Older ASICs lacks such capabilities resulting routers more software dependable and may be limited in table size capabilities.

Solution: Can whitebox Switch help?

Disaggregation or whitebox is the best solutions for this problem. Buy the choice of hardware herein router/switch from your preferred vendor and select software or Network Operating System (NOS). The benefit of whitebox or disaggregation for that matter allows you to select best of the breed merchant silicon and buy those in your terms with a price point you can afford. Result you get the best of both world: hardware and software. For the bigger IPv4 table size, Broadcom® Qumran-MX™ silicon with BCM52311™ Knowledge-Based Processor (KBP) [TCAM] provides you optimal choice for upto 1 million IPv4 routes. There are cases where upto 1.2 million routes are possible in such system.

Figure 3. Edge Router Whitebox based on Broadcom® Qumran-MX.

A number of hardware vendors are currently offering Qumran-MX based platform with industry proven NOS from companies such as IP Infusion.  As depicted in the figure above, Merchant Silicon herein Broadcom® Qumran-MX™ is connected through an internal bus (known as ELK bus) to external TCAM which provides further capabilities for lookups.

However, it is also important to select appropriate software vendor that has optimize such boxes and provides optimal route capabilities of more than 768K to facilitate your upgrade or help in your preparation for 768K day.

IPInfusion’s OcNOS™ is tested with a number of Hardware vendors providing you a wide slections, please ensure you select appropriate Qumran-MX based hardware with External TCAM to achieve upto 1 Million route. If you are interested about OcNOS and how it can solve your 768K day, you may visit their website at https://www.ipinfusion.com .

However, please make sure you ask each vendor to provide you with test report or atleast enough data to make educated decision.

 

Reference

[1] Some internet outages predicted for the coming month as ‘768k Day’ approaches. Available at https://www.zdnet.com/article/some-internet-outages-predicted-for-the-coming-month-as-768k-day-approaches/.

[2] Australian Internet Users Face Looming ‘768k Day’. Available at https://www.cso.com.au/mediareleases/34669/australian-internet-users-face-looming-768k-day/

[3] Australian Internet Users Face Looming ‘768k Day’. https://www.computerworld.com.au/mediareleases/34669/australian-internet-users-face-looming-768k-day/.

[4] Thousand Eyes. What is 768K Day, and Will It Cause Internet Outages? Available at https://blog.thousandeyes.com/what-is-768k-day/.

[5] Internet outage tag at twitter. Available at https://twitter.com/search?q=internet%20outage&src=tyah

[6] CIDR, 2019. CIDR report for June 11, 2019. Available at https://www.cidr-report.org/as2.0/

Network Transformation is in the offing: Does disaggregation model help?

One of the fascinating aspects in my job is understanding how some of world’s largest networks will deploy Open networking platforms. It is this continuous learning that keeps me passionate about what I do.  When I first use “open networking” term, people wondered what it may entails.  Some thought of it is all about using “open source” software and the other thought of it as another connotation of confusing technical term that simply add to plethora of jargons.  Interestingly, open networking is none of the said. It is about disaggregating hardware and software in a way that allows relatively easier integration of each sub systems. In doing so, the model innately provides three distinct benefits to customers:

  1. Eliminates Vendor lock-in: There is no need to be enchained by vendor or vendors. With state of the art merchant silicon available off the shelf, benefits of technology can be realized without being locked into a specific vendor device.
  2. Cost Reduction: As more and more vendors join the marketplace, price of hardware and software becomes suddenly more affordable. This innately reduces CAPEX and OPEX for customers.
  3. Choice of HW and Software: The disaggregation model in which open networking conception is born innately provides a choice for customers on hardware and software.

 

Figure 1‑1. Open Networking: Disaggregation model.

Central to open networking is the disaggregation model that allows hardware and software vendors to deliver their capability on what they do best. For example, hardware vendor can bring best breed of technology from modularization of their platform to make use of state of the art market silicons and on the other hand, software vendors to offer best NOS (Network Operating Systems) or software means for data plane and control plane separation.  In doing so, the marketplace has created an apparently robust eco-chain. Some will disagree here but the fact remains that the eco-chain is improving and some parts of it is more realistic now then perhaps a year before. To be more specific, I have elaborated the eco-chain in the figure below. Rather than referring it to open networking eco-chain, I denoted it as “SDN eco-chain” since the central to this initiative is about making network more programmable and service aware.

 

Figure 1‑2. Typical SDN Eco-chain.

The programmability is the core element of SDN (Software Defined Networking) for which pundits advocated the decoupling of control and data plane. However, programmability can be done in many ways: through ZTP (Zero touch provisioning), APIs (e.g. Netconf) and OpenFlow protocol (a more conventional notion of decoupling).

To my experience, majority of network deployments thus far refrained from purely decoupling model but there are exceptions and distinct differences in deployment. For example, FBOSS that runs as agent daemon uses thrift API instead of openflow to communicate route information with controller. Details of this opensource decoupling software is available at https://github.com/facebook/fboss .

Service Providers are more interested about CAPEX and OPEX reduction with choice and flexibility that open networking offers. For example, a provider edge switch that offers flexibility of advance MPLS networks with APIs for model based network configuration and monitoring could best fit their immediate need of transforming networks as oppose to overhauling the network in one go using decoupling of data and control plane. It is in this phased transformation, open networking offers best value but it is not confined by the said as it is encompassing to traditional networks and decoupling of data and control plane.

I conjectured even a bigger picture for “Open networking” as it is the fundamental conduit of future network transformation. For it, disaggregation model is paving the way for a more sophisticated network transformation to occur. A vision towards that end would be “Intelligent Networks” one that is provisionable in click, intent-based and self-healed, to name few attributions.

Figure 1‑3. Networking Technology Trend.

Affordability, openness and choices are big factors in network transformation especially for service providers since disaggregation model is giving them ability to deliver more service at a fraction of cost that would otherwise be. Moreover, open networking is offering best of both world, the “traditional” and “NextGen”. For example, in a given network deployment scenario customer may prefer to choose their own NOS depending upon southbound API requirements and L2/L3 switching products. A NOS that combines both APIs and decoupling protocols such OpenFlow is known hybrid NOS. The goal here would be two folds, first allow traditional network protocols such BGP to be run at NOS level forming traditional routing architecture and secondly, network topology is passed to Openstack for visibility, orchestration and management. The following picture depicts how such nextgen network abstraction can coexist with traditional networks.

Figure 1‑4. Cloudification of Network.

Telemetry at each network device can be obtained through Broadview™, an opensource software contributed by Broadcom®. Such deployment not only allows physical networks to be virtualized and cloudified, it is done so without compromising backward compatibility of physical network and the traditional network routing architecture therein. The depiction herein is forward looking and implementation of such network is at the discretion of its designer and availability of suitable hybrid NOS. However, this notion of network cloudification is possible only due to advances in open networking and the eco-chain that supports it.

This article is provided herein for educational purpose only implying the importance of “Open Networking” and how disaggregation model helps in realizing NextGen network transformation.