Showing posts with label standards. Show all posts
Showing posts with label standards. Show all posts

Wednesday, October 1, 2025

A Manifesto for Data Realism

The contemporary discourse on data governance has been compromised by Data Idealism which approaches data as primarily a social and techno-legal artifact. There are islands of data idealism, such as data ethics, "Free Flow with Trust", data decolonisation, data feminism, ethical AI development and others which basically suggest that much of the social, political, and economic consequences of our digital age can be managed through the mechanisms of transparency, fairness, and ethical alignment. This is a case of structural blindness. Here, I propose Data Realism as a necessary corrective, requiring a shift in focus from the (laudable) ideas of computational equity to the more concrete realities of infrastructural ownership, standardization leverage, and strategic capacity to manage this non-fungible asset.

This requires recognising the following Five key tenets of Data Realism:

1) The world exists and data are our only contact with it. 

To deny this is to abandon the epistemic project - data may be imperfect, shaped, mediated — but contact with the world nonetheless. Once we accept the fundamental role of data as interfaces to reality, Data Realism demands that broad, cost-effective access to data be made possible. The goal is to create and maximize the utility of data commons while minimising systemic risk. Today, AI companies and developers need clarity on what data they can use and how. A facilitative sourcing framework will remove the constant threat of litigation, allowing development teams to focus on quality and performance of their models rather than worry about legal risk management. The current public data bottleneck stifles competition in AI.

This focus on "making things work" means Data Realism advocates for policies that legalise the collection of publicly available data. The current legal ambiguity and ethicist shaming cripples startups. Surely, there have to be clear technical standards for scraping (rate limits, robots.txt adherence, mandatory anonymisation, exclusion of sensitive and non-essential data etc), but by lowering the cost of basic data access and creating data commons, states can forces AI companies to compete on superior modeling, contextual application, and algorithmic innovation — rather than on who is the biggest and baddest proprietary data hoarder.

In lieu of this, more public and private investment in curated, contextual public datasets are needed. These vetted datasets can lower the initial data sourcing cost for startups and create a standardized benchmark for model development, replacing expensive, ad-hoc, and legally risky scraping efforts. Regulatory policy here must mandate data sharing or standardized APIs for essential public-interest data held by natural monopolies by incentivising voluntary contribution of anonymised, high-quality datasets to open-source commons. Further if data are our contact with the world, an over-reliance on specific metrics can warp signals, so data realism also demands holistic reality capture that incorporates qualitative insights and a plurality of indicators.

2) Data exist with the world. 

This implies that data production is situated and filtered through the environment - data are not an abstraction but a critical resource which come from somewhere, are made by someone, and are shaped by instruments, protocols, and power. Data not only represent but also enact realities - especially as more and more information systems are automated - data shape public discourse, inform policy, and modulate real human and machine behavior. The materiality of data infrastructures is far from ephemeral, data are stored, circulated, and maintained by physical systems that leave significant ecological and geopolitical footprint. Since every interaction leaves a trace - Data Realism demands we acknowledge that data is not simply "collected" but its genesis and production is infrastructured. When schools of data idealism focus on moral arguments about data, they also accept the infrastructural dominance of incumbent hegemons (and their ethical priorities) as an unchangeable premise, seeking to ameliorate the prevailing system of power rather than challenging its foundations. 

Therefore, a successful Data Realist state must foster a "permanent view of politics" required to integrate the trajectory of global technological developments into its own strategic calculus, prioritize the development of sovereign technical and even management standards surrounding data, and explicitly link digital industrial geographies to national security goals. Just as the world has winners and losers, the digital society has data powers and data provinces.

3) The world leaks through.

Data realism is not a defense of surveillance, dashboards, spreadsheets, or technocratic governance. It is a defense of reality as something external to human discourse and design — something that can resist, surprise, and falsify our models. To that end, Data Realism rejects two dominant trends:

   Naive Empiricism — the idea that data “speak for themselves,” that numbers are neutral, measurement is innocent. This view fails to account for context, biases, or interpretation.

   Radical Constructivism — the view that data are nothing but power-laden constructs, shaped entirely by ideology, narrative, and positionality. This view erases the world and collapses epistemology into politics, often for sake of it.

A realist stance rejects both the blind faith in datafication and the nihilism of pure relativism. Data realism does not deny context, ideology, or structure. It insists that, even through those, the world leaks through. Consider a temperature reading, a mortality rate, or a vote count. These are not just narratives. They constrain us. To treat data as real is to take them seriously — not as final truth, but as our provisional contacts with the world. It is to ask what all this shows and means, not just who made it and why. Data therefore must be analyzed without idealization, where a commitment to the hard facts of data, even if inconvenient or ugly (e.g., showing inequality, corruption) is necessary and statistical gaslighting civilisationally poisonous. Data can be manipulated. But manipulation presupposes a baseline that can be distorted. Falsifying a vote count still depends on the idea of a real vote count. Censoring mortality rates still implies that there were deaths. To lie with data is to admit that truth matters, because at some level data are a non-negotiable reality that exist and operate independent of our political beliefs and moral aspirations. This means we must apply the highest scrutiny to the data used to train and test our systems, human or artificial. 

4) Data drive agency in the world.

Data is not an end but an index of industrial, military, and academic capacity. A state’s true data capacity is measured not by the size of its population’s data footprint, but by its independent ability to standardize, store, and compute that data without reliance on external supply chains or governance frameworks. This requires a systemic integration of military, academic, and industrial objectives — a union of science with industry that treats digital technical standards as global public goods that must be wielded strategically, and not just consumed passively. Contrary to idealistic claims, national security is the ultimate policy engine driving data governance decisions at the level of states, with privacy and ethics serving as secondary and often negotiable constraints. The data policies of the hegemonic and emergent powers are fundamentally rooted in securing technological advantage. The task therefore is not to eliminate dependency through isolation, but to gain the necessary leverage in data systems to shape the rules of its game.

Idealistic data policies are politically naive because they assume consent and cooperation in an anarchic system. Consider the G7's DFFT narrative, for instance, which is often projected as a universal good, but is mostly an elegant rhetoric of "free flow with trust" that uses an abstract legal promise as a mask to hide the concrete realities of global political controls left unacknowledged. A Data Realist state, therefore, must subject all policies to a simple but rigorous test: Does this policy measurably increase sovereign capacity and reduce structural dependency, or does it merely achieve moral compliance with the powers that be? Data Realism thus demands a meaningful shift from the judicial-police state (focused on making and enforcing laws) to the structural state (focused on building and owning digital capability and metapolicy spaces). 

5) Navigating the world with data requires pragmatism. 

Data realism is commitment to practical ethics, not idealisms. It eschews notions of diagonally opposite left/right systems. It is a philosophy of effective technological acceleration and not of technological pessimism. Practical ethics require direct and immediate confrontation with ethical necessities of data flows - make the methods of collection, cleaning, modeling, and interpretation as transparent as possible to those who are affected by the resulting decisions - but beyond a right to audit and redress, data stewards should not have to project desires of how the world or social contracts ought to be into their data pipelines. Data ownership is the ownership of Truth, and thus carries a responsibility to protect and de-risk the data in their care, and if required, transfer that ownership for systemic continuity. The primary task of governance here is thus not to make data and systems ethical, but to confront and master the measurable structural facts of computation, ownership, and capability. 

As almost everyone knows, the world and its governments are secretly run by accountants. This implies that data should go through continuous assessment for depreciation or appreciation. Once you put a number on the decay or change in data's subjective value due to context shifts and bias development, it will better incentivise the appropriate and timely flow of organisational resources to update and responsibly address the state of its data pipelines - as necessary to hold on to the realism in data. The financial incentives for better data governance and lifecycle management, reducing long-term infrastructural and technical debts, should thus be made explicit and immediate to the accounting elites. 

To conclude, as AI and automated systems gain more and more leeway into human affairs, this manifesto calls for embracing Data Realism, a philosophy anchored in the undeniable existence of the world and data's fundamental role as our provisional contact with it. It is a mandate to recognise that in a geopolitically volatile world, reliance on external, proprietary, or geographically constrained data sources can be a major systemic vulnerability - and argues for a strategic shift toward data resilience, reliability and sovereignty to guarantee uninterrupted operational continuity of our digital lives regardless of external regulatory or political pressures.

Sunday, May 18, 2025

Some Cheeni Technicalities


Recently, for an academic workshop I had to submit a brief on China as a latecomer in global technology standardisation. Notwithstanding the Napoleonic assessments of the waking of a sleeping giant, I argue that the Chinese “latecoming” in global technology standardisation has been more of a “coming of age” instead. As Kurt Campbell and Rush Doshi note, presently China has twice the manufacturing capacity of the US, is churning out more active technology patents and highly-cited papers than the US, not to mention of producing 80 percent of consumer drones worldwide, and boasts of a shipbuilding capacity 200 times larger than the US. This structural transformation has been decades in the making, and should a single word be chosen to describe the Chinese strategic thinking in light of its “latecoming”, we may concur that word to be Patience.  

In an earlier blog, I had noted the sharp rise in Chinese contributions to internet governance from 2008 onwards, eventually coming head to head with a slowly declining US in the making of a bipolar world. Let us unpack the structural factors that led to that sharp but stable rise over the years. 

First of all, around mid-2000s, China began a twofold process of global influence. One aspect of it was to synergies its military and economic policy under the slogan of Zìzhǔ Chuàngxīn which could be translated either as indigenous or independent Innovation. And the second was Zǒu Chūqù which could be translated as Go Out (as in global). By the end of the 2000s, China had thus tried to reform its domestic industry and enabled them to capture global markets. Consider that in 2008, 75 percent of Huawei’s total revenue came from outside China, demonstrating the successful internationalisation of Chinese industry. Ostensibly, China had to undertake a rapid drive towards international technical standardization to support its global trades, especially from ICT corporations like ZTE and Huawei. 

Simultaneously, China also completely revamped its domestic R&D environment. The Chinese state actively enabled people in its technology companies and academia to develop strategic cooperations in national interest and participate in various multistakeholder technology forums. A lot of this also came from requiring PhD outputs to contribute to standards or open-source type of ecosystems, as well as the effective use of diaspora networks. This has significantly driven academic actors like Tsinghua University, as well as state actors like China Telecom and China Mobile, towards actively shaping the outputs of transnational technology forums. We must note here that in this China’s own digital transformation required a standards backbone which these actors found the opportunity to build themselves – leading to such entities producing successful global standards in areas like IPv6 transitioning or the multilingual internationalization of domain names.     

China has also been strategic about the use of its territory to engage with global talents. In November 2010, IETF’s 79th meeting was held in Beijing, which alone may have cause a significant spike in Chinese technical publications during 2010. In comparison, other “latecomers” have not made such explicit attempts to engage with global communities of experts. India, for example, has not yet hosted any major IETF meetings (barring some small on-boarding sessions in universities) despite a large IT sector and several digital-industrial geographies like Bangalore and Hyderabad. What further stands out is that the Chinese spike of 2010 has sustained and improved in the subsequent years, indicating that change emerged from strong structural foundations. 

In all, China vies to become the new global hegemon. Ostensibly, a global hegemon has to provide global public goods, and open digital technical standards are a prime variety such global public goods. DeepSeek’s surprising open-sourcing too had followed the same logic and even technical goals (relating AGI frontiers). The fluctuations in US’ determination to keep providing such key public goods in digital technologies, as has been evident in the recent back-and-forth over the the financing of CVE databases, or in the quite migration of RISC-V microprocessor governance to Switzerland, further creates a vacuum which other entrepreneurial states arriving “late” into this onsetting global bipolar disorder may try to fulfill through their own structural and productive capacities - if they know how to nurture and wield those capacities.  


Saturday, February 1, 2025

Infrastructuring AI in the Postcolony

The root of all evil is a premature policy optimization.

In the few decades of the history of computing in India there have been some spurts of proactivity, but by and large other than a large bodyshop in IT, we have been comfortable in a mode largely reactive to global developments, with an industry making negligible contribution to computing technologies. Just after liberalisation, C. R. Subramanian had undertaken a long-view analysis of the matter in his "India and The Computer" - with many of his suggestions ringing true to folks in AI today. Let me therefore reconstruct below the difficulties that were noted over 30 years ago for large scale "harmonious development" (not support and maintainance) of computing and software products in India:

1. The domestic markets in India are too small. To make a mark, we've to address the international high-technology demand, which means society being at the cutting-edge of technology and R&D. Considering quick internet based distribution, open-sourcing and market standardisation, any catch-up oriented strategy would have to fight against strong network effects as well. This will be very hard for domestic industry to do in presence of established foreign competition without some protections until they reach a certain stage of technological and business maturity. 

2. Subramanian clearly recommends 'standardisation first' to be one possible technology strategy, but he also says, "...there is a total lack of concern for this at the highest levels of the administration. Political support has not even been sought for this vital step." Based on my own interactions with the Bureau of Indian Standards and other IT sector actors, 30 years later Subramanian's characterisation of embedding political vision in India's technical standardisation strategy has not changed much. Stock policy formats of "one nation one..." lack compatibility with global standardization environment - one way to address this would be to dissociate IT standards from BIS (which comes under Consumer Affairs) and give it directly to MeitY and technology industry stakeholders bodies who have skin in the game.

3. He says, "the private sector in India does not traditionally invest in high-risk, new technology areas." One of the main reasons for this, perhaps owing to our socialist traditions, is that the industry actors (at least the smaller ones) are not seen with much trust and respect. Hans-Peter Brunner notes that 1984 was perhaps the first time since independence when domestic industry actors in India were given on paper a “respectful place”. The little industry guy is not really accorded the respect of a 'partner' in public-private interactions, not to mention of the babu expectations of profound subservience. Resultingly, the industry has stuck to relatively risk averse and by-the-book activities, having neither the financial incentives nor the resourcefulness for high-technology risks/innovation. 

4. Infrastructural ownership, he says is a "must from the point of view of national security and related developments in the space and defense areas. To allow the World Bank a say in the matter is inviting foreign interference in domestic technology aspirations related to self-reliance." Considering our situation in the semiconductor supply chains and an unwritten policy to export brains (the human infrastructure), and now also with data and platforms, who really has digital infrastructural ownership in India?

5. He notes that the "inconsistency in the various policies — fiscal, licensing, technology development and technology import - prevent harmonious development". The natural evangelist of technology in government anywhere is the military, which faces a policy conundrum in India between satisfying its immediate requirements for finished products vs sticking it out for the development of the "indigenous" industry - unless in view of Atmanirbharta it could take a more strategic role in domestic technology sector. I may remind the reader here that the true union of science with industry happened only in the WWII, which in the US resulted into a cultural osmosis where the military picked up management perspectives and businesses picked up military ones.            

But how soon can these difficulties be overcome? In a world where technology infrastructures are often intermingled with social and political imaginations, a mere exorcism by diplomacy and public relations cannot address India's currently vacuous strategy for technological leadership. Our national security as well as digital technology governance being run predominantly through a policing lense has not halped our cause either. Firstly, I must say that it is not the technical but the management standards in our organisations (including govt and startups) which need to catch up with the engineering talent.   

Thence, the AI hype can be appealing to broken institutions - and can lead quickly into a procurement mentality where run-of-the-mill, off-the-shelf analytics and interaction products (even LLMs wrapped into brands) are acquired in the name of promoting and enabling the AI research ecosystem. That is like getting a website made to enable the development of future internet. Shoppers have to shop, but without deeply trusting your own people, letting them not go, and persevering with them on novel ideas, the technological catch-up would not end very soon. Perhaps it is also time for us to deploy legally a broader definition of CSR finance to include certain kinds of R&D as well.

Consider also, that organisational innovations are needed to integrate technology as a doctrinal component at the highest levels. We do not need an "AI policy" as much as we need to formulate a multi-stakeholder industrial policy that integrates and addresses different sector-specific demands (and expectations) of deploying automation, data processing and cyber security. This is a must for national competitiveness, security and defense, and cannot simply be exorcised by witty diplomats.  

Mbembe writes on his experience in postcolonial Africa, that often the actual development of infrastructure has more to do with access to government contracts and rewarding patron-client networks, and not its technical function or nuances - and some note that "this is why roads disappear, factories are built but never operated and bridges go to nowhere"- as we might sometimes relate as well. Now imagine a super-intelligent AI being developed and deployed in the purview of a similar organisational paradigm. If that is the sovereign national discretion, so be it, but it might be useful to make available trustworthy infrastructural components - standardised and open building blocks, both logical and physical, that can be integrated (and disintegrated) at base/local levels as required by anyone - as the local is not very far from the global here - for pure technology competition, for all practical purposes, implies an overall escalation of force with non-lethal effects.  


Monday, May 27, 2024

Technology and Security Competitions

There circulate these days a variety of rumors of technological destruction - from an impending AI doom to a US-China technological decoupling and subsequent showdowns, or even an escalation of "small wars" into a litany of CBRN scenarios - for all we know (do we ever?), the internet could well be full of competing social botnets programming us for our imminent futures. Nevertheless, these rumors invite a closer look at the nature of technology development, coupling and competition before we could mourn (gleefully?) upon the demise of this brave new world of ours. In this digitally inter-connected world where "there is no there and we're all here", geography still remains a fundamental reality. Westphalian states are the dominant security providers over specific geographies. There are many competing states, perennially uncertain of each others' intentions and desirous of each others' resources and geographic control. They must also maintain a moral upper-hand and legitimacy over competing security providers. This security competition is one of the the key troubles in contemporary governance of "cyberspace" where there exist competing multilateral vs multistakeholder logic and value systems. 

Let us therefore begin with the internet itself. The graph below depicting the nation-wise publication of technical documents at the Internet Engineering Task Force (IETF) underlines the nature of technological contest among "great powers". A Thucydides vibe between a rapidly rising power (China) and a gradually declining one (US) is quite naturally apparent here. Ostensibly, the cyberspace is a global common, contested but shared, and its protocols' and standards' development, historically contingent as is it, could not be used to make generalised remarks about the development of AI technology stack and its technical standards. For those are developed with a much greater regional impetus where locally clustered actors dominate markets and policy. Moreover, as AI and social bots become more pervasive, internet governance itself might have to integrate aspects of geographically contingent platform and API governance machinations.       

A Thucydides' Graph of Technology Competition [Source]

This geographic characterization of digital stack portends an incipient geopolitical logic in technology construction. To take a particular area, one could see how geography forces technology in the development of various national cyber security complexes. For example, the US, that has no mortal enemies at its borders and is fairly isolated by the oceans, but has to fight wars all over the world. It has, as a result, leaned quite heavily on developing global communications and surveillance networks, superiority in air and electronic spectrum, and the enabling cyber security partnerships such as Five Eyes. Furthermore, one may argue that the present internet norms and architecture in itself are a significant tool at the service of the liberal international order.

The Chinese cyber security complex, on the other hand, has leaned a lot more towards domestic control, industrial espionage, management and expansion of territory, and its ambitions over pacific waters - where it comes into direct conflict with the US hegemony that'll continue to shape its digital-technical goals. One must note here that the construction of state's cyber security complex entails all three properties of technology - namely technique, equipment, and organisation - the unique nature of which emerges from the geo-strategic forces underlying these. 

One of the best example of this geographically contingent interplay of social organisation, technique, and equipment is Israel. Its location and initial conditions forced it to shed all the useless pomp and hierarchy of military-technical organisations and adopt a ruthless functionalism instead, producing in effect a world-class cyber security complex without the accompanying burden of quasi-victorian bureaucracies. In fact, Russia's embrace of "hacker culture" and asymmetric cyber capabilities with respect to US and Europe must also be seen within the context of the collapse of Soviet geographic and economic vision, and not to mention the abject failure of Soviet's symmetrical competitive strategy with US in early internet development.    

We, not to leave ourselves out, have had two main geographic adversaries - Pakistan, and China. However, early on our state managers pivoted the Indian strategic thinking and discourse around Pakistan, not China, and stuck to it. Pivoting our strength and capability building around the smaller adversary was certainly easier and also suited the political and professional incentives of powers that be, given long and painful historical damages. However, this long-held strategic benchmark of a useless and weaker enemy produced a psychological and technical backwardness in our society - we chose to import, not build, our own military-security stack, including even cyber security software. In fact, it was a third-party cyber threat observer (who we also tried importing then, and is now famous for its Pegasus investigations) that had initially flagged the sweeping extent of Chinese botnets in India to the notice of the public and government.   

The geo-strategic foundations of technology development suggest that as long as our polity remains shy and inertial about the deeply geographic and military-technical nature of our competition with China, our cyber security complex and technology organisation too will continue to reflect that institutional ambiguity. There are three key broad lessons here requiring "radical acceptance" by policymakers that the geo-strategic foundations of technology development hold: 

A) Technocratic rationality dominates ethical rationality in hyper-competitive arenas. 

The development of dominant cyber security clusters in Tel Aviv, Washington and San Francisco indicates how closely tied together innovation is with knowledge spillover and a hyper-competitive social ecosystem. Not only is there a great mobility of high-end expertise across public-private organisations within these geographies, it also corresponds highly to the military-security requirements of their states. Hyper-competitive games cannot be played with ethical or ideological instruments. A case in point can be that present AI systems need robustness, safety, and energy efficiency - as per technocratic rationality - but if governments prioritize corporate DEI policies instead and direct finite resources into inter-governmental virtue signaling games, they may win some validation but will lose the broader security competition itself.  

B) Global technical standardization is beyond conventional competencies of governments. 

The rapid rise of China in the Thucydides' graph above draws significantly from contributions of companies like Huawei, Baidu, and Tencent, along with actors like China Telecom. This happened post a series of reforms in early 2010s giving more leeway and independence to such actors in the global internet governance. Thus, transnational technical standardization requires a whole-of-society anti-Westphalian approach to address certain gaps in states' technical expertise and governance capacities. Having to navigate this conflagrating techno-security competition between US and China, we need serious structural reforms now across the state and industry to dramatically drive the trajectory of our own line, and certainly not more of the same thing.  

C) Technology artifacts are not technology. 

Fundamentally, technology is information (as in know-how). One may further say that technology is strategic information. On their own, societies acquire this knowledge under considerable security and civilisational stressors, wars being a prominent one. Since it is widely understood that high ambitions often have to grapple with constrained timelines and low budgets, it is tempting for bureaucracies to buy cool stuff and say that they've acquired technology. Yet, this is eventually hogwash accompanied with an overbearing servitization component. The Americans and Israelis got this information in process of navigating warfare, and the Chinese by stealing intellectual property instead.     

This discussion also underlines a key change in national polities post-WWII (because of the introduction of nuclear weapons and a corresponding military-scientific elite in decision-making processes) that in addition to the military and geopolitical planning, states also have to integrate global technological developments in their strategic calculus. With strong AI around the corner and a possibility to remake the internet (for good?), this would require considerable expertise outside the usual skills possessed by politicians and their babus, and additionally a permanent view of politics beyond electoral vicissitudes. Technology, by transforming our needs, environment and actual possibilities, slowly shapes its own operating environment. It has a life, and now a mind, of its own, throwing the transnational digital governance in practice into an uncomfortable mix with the contraptions of the administrative state. Where does our geo-strategic imperative take us at this juncture?

Monday, September 11, 2023

Reconciling AI Governance and Cybersecurity

  

Recently, Sam Altman had been touring the world attempting (perhaps) a regulatory capture of global AI developments. No wonder that OpenAI does not like open-sourced AI, at all. Nevertheless, this post isn’t about AI development, but its security and standardization challenges. Generally, an advanced cyber threat environment, as well as the defenders’ cyber situation awareness and response capabilities (henceforth referred together as ‘security automation’ capabilities) are both overwhelmingly driven by automation and AI systems. To get an idea, take something as simple as checking and answering your Gmail today, and enumerate the layers of AI and automation that can constitute securing and orchestrating that simple activity. 

Thus, all organisations of a noticeable size and complexity have to rely on security automation systems to effect their cybersecurity policies. What is often overlooked is that there also exist some cybersecurity “metapolicies” that enable the implementation of these security automation systems. These may include the automated threat data exchange mechanisms, the underlying attribution conventions and knowledge production/management systems. All these enable a detection and response posture often referred to by marketers and lawyers as “active defense” or “proactive cybersecurity”. However, if you pick up any national cybersecurity policy, you’d be hard pressed to find anything on these metapolicies – because they are often implicit, brought into national implementations largely by influence and imitation (i.e. network effects) and not so much by formal or strategic deliberations.

These security automation metapolicies are important to AI governance and security because in the end, all these AI systems, whether completely digital or cyber-physical, exist within the broader cybersecurity and strategic matrix. And we need to be asking whether retrofitting the prevalent automation metapolicies would serve well for the future of AI or not. 

Avoiding Path Dependency 

Given a tendency towards path dependency in automated information systems, what has worked alright so far is getting further entrenched into the newer and adjunct areas of security automation, like the intelligent/connected vehicle ecosystem. Further, the developments in the security of software-on-wheels are being readily co-opted across a variety of complex automotive systems, from fully-digitised tanks that hold the promise of decreased crew-size and increased lethality to standards for automated fleet security management and drone transportation systems. Consequently, there is rise in vehicle SOCs (Security Operations Centers) that operate on the lines of cybersecurity SOCs and use similar data exchange mechanisms, borrowing the same implementations of security automation and information distribution. That would be perfectly fine if the existing means were good enough to blindly retrofit into the emerging threat environment. But they are far from it. 

For example, most of cybersecurity threat data exchanges make use of the Traffic Light Protocol (TLP), however TLP itself is only a classification of information – its execution, and any encryption regimes to restrict distribution as intended, is left to the designers of security automation systems. Thus there is need for not just more fine-grained and richer controls over data sharing with fully or partially automated systems but also ensuring compliance with it. Much of threat communication policies like the TLP are akin to the infamous Tallinn manual, in the way that these are almost an expression of opinions that cybersecurity vendors may consider implementing, or may not. It gets more problematic when threat data standards are expected to cover automated detection and response (as is the case with automotive and industrial automation) – and may or may not have integrated an appropriate data security and exchange policy for lack of any compliance requirements to do so.

Another example of inconsistent metapolicies, out of numerous others, can be found in the recent rise of language generation systems and conversational AI agents. The thing is that not all conversational agents are ChatGPT-esque large neural networks. Most of these have been in deployment for decades as a rules-based, task-specific language generation programs. Having a “common operating picture” by dialogue modeling and graph-based representation of context between such programs (as an organisation operating in multiple domains/theaters could require) was an ongoing challenge before the world stumbled upon “attention is all you need”. So now basically we have a mammoth of legacy IT infrastructure in human-machine interface and a multi-modal AI automation paradigm that challenges it. Organisations undergoing “digital transformation” not only have to avoid inheriting the legacy technical debts but also consider the resources and organisational requirements for efficiently operating an AI-centric delivery model. Understandably, some organisations (including governments) may not want a complete transformation right away. Lacking standardised data and context exchange between the emerging and the legacy automated systems, many users are likely to continue with a paradigm they are most familiar with, not the one which is most revolutionary. 

In fact much of cybersecurity today hinges on these timely data exchanges and automated orchestration, and thus these underlying information standards become absolutely critical to modern (post-industrial) societies and the governance of cyber-physical systems. Yet, instead of formulating or harmonising the knowledge production metapolicies needed to govern AI security in a hyper-connected and transnational threat environment – we seem to be falling into the doomer traps of existential deliverance and unending uncanny valleys. That said, one of the primary reasons for the lack of compliance and a chaotic standards development scenario in security data production is the lack of a primary governance agent. 

The Governance of (Cyber) Security Information

Present automation-centric cyber threat information sharing standards generally follow a multistakeholder governance model. That means these follow a fundamentally bottom-up life-cycle approach, i.e. a cybersecurity information standard is developed and is then pushed “upward” for cross-standardisation with ITU and ISO. This upward mobility of technical standards is not easy. The Structured Threat Information eXpression (STIX), which is perhaps the de facto industry standard now for transmitting machine-readable Cyber Threat Intelligence (CTI), is still awaiting approval from ITU.  Not that its really needed, because the way global governance in technology is structured, it is led by industry and not nations. The G7 have gone to the extent of formalising this, and some even blocking any diplomatic efforts towards a different set of norms. 

This works well for those nation-states which have the requisite structural and productive capacities within their public-private technology partnerships. Consequently, the global governance of cyber technology standards becomes a reflection of the global order. Sans the naming of cyber threat actors, this still had been relatively objective by nature so far. But it is no longer true with the integration of online disinformation into offensive cyber operations and national cybersecurity policies – not only the conventional information standards can run into semantic conflicts, newer value-driven standards over information environment are also popping up. Since the production and sharing of automation-driven social/political threat indicators can be shaped by and affect political preferences, as the threats of AI generated information and social botnets rises, the cybersecurity threat information standards also slide from a sufficiently objective to a more subjective posture. And states can do little to reconfigure this present system because the politics of cybersecurity standards has been deeply intertwined with their market-led multistakeholder development.

Cyber threat attributions are a good case in point. MITRE began as a DARPA contractor, and today serves as an industry wide de-facto knowledge base for computer network threats and vulnerabilities. Of the Advanced Persistent Threat groups listed in MITRE ATT&CK, close to 1/3 of cyber threats are Chinese, another ~1/3 come from Russia/Korea/Middle-East/India/South-America etc, and the remaining ~1/3 (which contain the most sophisticated TTPs, the largest share of zero-day exploitation, and a geopolitically aligned targeting) remain unattributed. We’ll not speculate here but an abductive reasoning about the unattributed threat cluster may leave readers with some ideas about the preferences and politics of global CTI production. 

A fact of life is that in cyberspace power-seeking states have been playing the role of governance actors and sophisticated offenders at the same time, so this market led multistakeholderism has worked out well for their operational logic – promulgating a global politics of interoperability. But it is bad for the production of cyber threat knowledge and security automation itself, which sometimes can get quite biased and politically motivated over the internet. Society has walked this path long enough to not even think of it as a problem when moving into a world surrounded by increasingly autonomous systems.  

A Way Forward 

With the social AI risks looming larger, states intending to implement a defensible cybersecurity automation posture today might have to navigate a high signal-to-noise ratio in cybersecurity threat information, multiple CTI vendors and metapolicies, as well as constant pressures from industry and international organisations about “AI ethics” and “cyber norms” (we’ll not venture into a discussion of “whose ethics?” here). This chaos, as we noted, is an outcome of the design of bottom-up approaches. However, top-down approaches can lack the flexibility and agility of bottom-up approaches. For this reason, it is necessary to integrate the best of multistakeholderism with the best of multilateralism. 

That would mean rationalising the present bottom-up setup of information standards under a multilateral vision and framework. Because while we do want to avoid partisan threat data production, we also want to make use of the disparate pool of industry expertise which requires coordination, resolution and steering. While some UN organs, like the ITU and UNDIR, play an important role in global cybersecurity metapolicies – they do not have the sort of top-down regulatory effect needed to govern malicious social AI over the internet, or implement any metapolicy controls over threat sharing for distributed autonomous platforms. Therefore, this integration of multistakeholderism with multilateralism needs to begin at the UNSC itself, or any other equivalent international security organisation. 

Not that this was unforeseen. When the first UN resolution was made assessing Information Technologies in 1998, particularly the internet, some countries had explicitly pointed out that these technologies will end up at odds with international security and stability, hinting at the required reforms at the highest levels of international security. Indeed, UNSC as an institution has not co-evolved well with digital technologies and the post-internet security reality. The unrestricted proliferation of state-affiliated APT operations is just but one example of its failure in regulating destabilising state activities. Moreover, while the council seems still stuck in a 1945 vision of strategic security, there is enough reason and evidence to relocate the idea of “state violence” in light of strategically deployed offensive cyber and AI capabilities. 

While overcoming the resilience of global order and their entrenched bureaucracies is not going to be easy, if reformed in its charter and composition, the council (or its replacement) could serve as a valuable institution to fill the void that emerges from the lack of a primary agent in guiding the security and governance standards driving security automation and AI applications in cyberspace. 

It Is The Process

At this point it is necessary that we call out certain misunderstandings. It seems that regulators have some ideas about governing “AI products”, at least the EU’s AI Act suggests the same. Here we must take a moment of silence and quiet to reflect on what is “AI” or “autonomous behaviour” – and it will soon dawn upon most of us that the present methods of certifying products may not be adequate for addressing adaptive systems rooted in continuous learning and exchanging data with the real world. What we’re trying to say is that the regulators perhaps need to seriously consider the pros-and-cons of a product centric vs a process centric approach to regulating AI. 

AI, at the end, is an outcome. It is the underlying processes and policies, from data engineering practices and model architectures to machine-to-machine information exchanges and optimisation mechanisms, where the focus of governance and standards needs to be, not at the outcome itself. Further, as software shifts from an object-oriented to an agent-oriented engineering paradigm, regulators need to start thinking about policy in terms of code and code in terms of policy – anything else will always leave a giant gap between intent and implementation. 

If the aforementioned chaos of today’s multistakeholder cybersecurity governance is anything to go by, for AI security and governance we need an evidence led (consider data that led to final CTI, and engaging with new types of technical evidence) threat data orchestration, runtime verification of AI-driven automation in cyber defense and security systems, clear non-partisan channels and standards for cyber threat information governance, and a multilateral consensus on the same. Focusing on the final AI product alone can leave much unaddressed and potentially partisan – as we see from the ecosystem of information metapolicies that drive security automation systems worldwide – hence we need to focus on better governing the underlying processes and policies that drive these systems and not the outcomes of those processes and policies.