Showing posts with label digital. Show all posts
Showing posts with label digital. Show all posts

Friday, May 22, 2026

Chasing Permanence: AI Replicas Against Entropy

 

In 2017, consistent with some traditions of preserving Tibetan lamas after their deaths, I had written of an advanced digital replica as a means to preserve the persona of His Holiness the 14th Dalai Lama. Since then, AI personas, replicas and digital twins, have come of age. In 2023, an actors’ union went on a strike demanding protection from generative AI applications. While AI replicas and man-machine symbiosis come with a number of societal hurdles that must be addressed (for instance, ethical and legal questions surrounding the ownership and rights of a digital consciousness) – I argue in this essay that AI replicas also present a unique and compelling pathway to preserve the core of humanity over an extended and potentially indefinite timeline across space, positing a post-anthropocentric reality of intelligence and agency in the universe. 

Some Grounding  

AI replicas build upon core foundations laid in early 1990s – one focusing on the instrumentation of a software robot and another on transferring human personalities onto a computer. The original notion of a “digital persona” stems from the latter, as introduced by Roger Clarke in 'Computer Matching and Digital Identity', who later reworked on the concept presented ‘the data model’ for use as a proxy for the individual whose information had constituted it. The original “software robots”, on the other hand, were developed as a fully implemented AI agent, whose sociality only began to seep into their architectural framing few years later. 

Clarke had envisioned the future of human digital replicas under passive and active modalities. His passive digital persona is merely a collection of data, personally identifiable, fitted into some data/information structures that represented some aspects of an individual’s reality. The active digital persona, on the other hand, is same data model with some actor-ness or agentic characteristics. The recent advances in human-computer interfacing, deep learning and generative AI capabilities are bringing this agent oriented paradigm to the center-stage today as a viable proxy for real human beings in their absence – thereby also driving a shift in data processing and management toward agent-orientation. This shift towards "agent-orientation" of information models also makes the preservation and orchestration of an "active persona" viable today. Yet, the dominant paradigms through which we digitally model humans remain fundamentally limited: capturing information but not agency, preserving states but not continuity of action. This gap is becoming increasingly consequential as AI systems begin to act on behalf of individuals across digital environments.

While terms such as personas, replicas, and twins are often used interchangeably, a key distinction lies in the directionality of their data connections. Digital twins operate through bidirectional flows, where physical systems continuously inform their digital counterparts and are, in turn, influenced by them. In contrast, most digital personas and replicas are primarily unidirectional, extracting and encoding information without directly shaping the underlying reality. But there is a gap between simulation and reality – which portends that the promise of AI replicas also carries some bio-logistical challenges. 

The Unique Promise of AI Replicas

AI replicas are not just static archives we are used to storing our knowledge and cultures into, but rather living, interactive repositories that can interpret, apply, and potentially even evolve knowledge. This puts AI replicas in a league of their own, completely different from other digital means of preservation. Even the contemporary developments in AI have primarily occurred as a mechanism to study, preserve and replicate the neuroscience of mind, and which in turn have led us to better AI models. Consider what human data really means here – identities, histories, and behaviors – embedded into a digital variant of the self which travels at light speeds. Of course, if your mind is running independently outside your skull in a computer in an environment you cannot even physically experience, it is a different agent, but one where your quirks and perspectives too live on. 

Two of our key points of departures from the conventional notion of digital twins in case of humans would be the wholeness and agency of the persona, and the absolute necessity of artifactual singularity. These are not trivial challenges and force structural abstractions upon the digital reality. A lot of internal bio-sensory data can be superfluous to the AI replica. However, same cannot be said about some of the other physiological characteristics such as blinking rates, muscle tension, pupil and skin responses – things which constitute our non-verbal communications. The holistic representation of personality including ethical, cultural, political, and emotional modeling of the human subject can also be fraught with some ethical, cultural, political and emotional effusions of various sections of society. 

The second sort of grand challenge, of maintaining artifactual singularity is even more daunting – if we do not get the engineering right, implications can be “soul crushing”. The challenge emerges from the very nature of digital objects – that they can be copied with little to no loss of information. Humans are unique, their DNA for the most part offers a unique fingerprint. We could make a digital copy of the human, as a passive or even an active digital persona. But it is absolutely critical for the portability of our social and cyber-physical realities that there remains only one unique “master” instance of a human being’s persona. Here the nature of the interface can also bring unique security and behavioral risks – but there are ways to address those, though none of them, being rooted in the cat and mouse game of securing cyberspaces, are close to perfect. Notwithstanding, through some mix of authentication, provenance about its creation, a global verification system, strict access controls and some guardrails on the active persona we could possibly engineer ourselves out of the precarious situations of facing multiple personality disorder in cyberspace.   

Stepping back a bit, note the temporal sweep of our existence. All this technological progress is very recent in human history and is rather asymmetrically distributed. Even today, stone-age implements (such as bullock-carts and mud-houses) co-exist with those being pursued at the frontiers of human civilization (such as AGI, space habitats, nuclear propulsion, gene editing). And this is just a few thousand years of history, and always at the brink of destruction. The fragility of our biological and institutional continuity creates a need for more durable forms of persistence. As civilizations now plan their travels from rock to rock across the universe, the case for merging the human identities into the AI systems becomes even stronger. 

Consider the case of someone like His Holiness the 14th Dalai Lama. The Tibetan sociocultural ethics, the world system, as well as His Holiness’ interpersonal relationships and understandings within that context would make his a very unique replica – a digital preservation like no other. If we localize that digital consciousness geographically, it begins to unravel even the aspects of a physical embodiment (although the Tibetan sage may himself retort that self consciousness is but an illusion to be transcended). Under the light of this particular instance, we may note clearly that the present preservation methods are not well equipped to handle the challenges of civilisational continuity – storing data is not the same thing as preserving an active persona – which sets up AI replicas as just in time to address such requirements. The one big question which needs to be pondered here is whether such an active persona will be autodidactic – meaning if the replica should continue to learn from and about the world on its own, long after the biological human whose information constituted the replica, may himself have succumbed to the inevitability of death. There are exceptional cases to concur upon this, where there can be clear consent and authority to steward the digital persona, even though it’d be infeasible at the scale of human populations. 

Ostensibly, a dynamic self-evolving digital consciousness is a beast unlike its biological counterpart. But if we were to try and replicate, or rather preserve even if not in function then in form, the human mind in silico – the technological affordances (or the lack of them) would force us to choose at present the level of abstraction we need for our digital brains. There is a gross organ level of the cognitive function, we could even go down into brain regions, may be even explore the neural microcircuitry, but then at the level of cellular and genetic simulacra we start hitting some serious hardware bottlenecks. Here I have purposefully left out matters such as 3D photorealism as problems tangential to the core human AI replica – although they are important in a world of active personas.   

Final Reflections 

Our discussion posits that a socio-technical movement in human computer interactions that goes much beyond the strictness of a simple human-tool relationship. This pursuit of permanence in-silico is not a rejection of our humanity but perhaps its ultimate expression to ensure our story endures long after we are gone from the universe. Even so, an evolving digital mind and consciousness need not be isomorphic to an animalist one, it is futile to search for homophily there and instead better to embrace faster a different ontology of socio-technical cognition. As many scholars have argued, this transference of reality onto software is only “an incremental evolution and not a radical departure”. Of course, there are aspects of our sociobiological reality which Turing machines cannot compute, such as non-algorithmic and non-deterministic events – which might also lend some weight to the idea that “consciousness” is a no go. But one may also, on the other hand, note here that this differentiation based on being in a continuous reality shatters upon consideration that this cherished continuous reality of ours is itself built upon a discrete one operating at much subtler levels. 

Wednesday, October 1, 2025

A Manifesto for Data Realism

The contemporary discourse on data governance has been compromised by Data Idealism which approaches data as primarily a social and techno-legal artifact. There are islands of data idealism, such as data ethics, "Free Flow with Trust", data decolonisation, data feminism, ethical AI development and others which basically suggest that much of the social, political, and economic consequences of our digital age can be managed through the mechanisms of transparency, fairness, and ethical alignment. This is a case of structural blindness. Here, I propose Data Realism as a necessary corrective, requiring a shift in focus from the (laudable) ideas of computational equity to the more concrete realities of infrastructural ownership, standardization leverage, and strategic capacity to manage this non-fungible asset.

This requires recognising the following Five key tenets of Data Realism:

1) The world exists and data are our only contact with it. 

To deny this is to abandon the epistemic project - data may be imperfect, shaped, mediated — but contact with the world nonetheless. Once we accept the fundamental role of data as interfaces to reality, Data Realism demands that broad, cost-effective access to data be made possible. The goal is to create and maximize the utility of data commons while minimising systemic risk. Today, AI companies and developers need clarity on what data they can use and how. A facilitative sourcing framework will remove the constant threat of litigation, allowing development teams to focus on quality and performance of their models rather than worry about legal risk management. The current public data bottleneck stifles competition in AI.

This focus on "making things work" means Data Realism advocates for policies that legalise the collection of publicly available data. The current legal ambiguity and ethicist shaming cripples startups. Surely, there have to be clear technical standards for scraping (rate limits, robots.txt adherence, mandatory anonymisation, exclusion of sensitive and non-essential data etc), but by lowering the cost of basic data access and creating data commons, states can forces AI companies to compete on superior modeling, contextual application, and algorithmic innovation — rather than on who is the biggest and baddest proprietary data hoarder.

In lieu of this, more public and private investment in curated, contextual public datasets are needed. These vetted datasets can lower the initial data sourcing cost for startups and create a standardized benchmark for model development, replacing expensive, ad-hoc, and legally risky scraping efforts. Regulatory policy here must mandate data sharing or standardized APIs for essential public-interest data held by natural monopolies by incentivising voluntary contribution of anonymised, high-quality datasets to open-source commons. Further if data are our contact with the world, an over-reliance on specific metrics can warp signals, so data realism also demands holistic reality capture that incorporates qualitative insights and a plurality of indicators.

2) Data exist with the world. 

This implies that data production is situated and filtered through the environment - data are not an abstraction but a critical resource which come from somewhere, are made by someone, and are shaped by instruments, protocols, and power. Data not only represent but also enact realities - especially as more and more information systems are automated - data shape public discourse, inform policy, and modulate real human and machine behavior. The materiality of data infrastructures is far from ephemeral, data are stored, circulated, and maintained by physical systems that leave significant ecological and geopolitical footprint. Since every interaction leaves a trace - Data Realism demands we acknowledge that data is not simply "collected" but its genesis and production is infrastructured. When schools of data idealism focus on moral arguments about data, they also accept the infrastructural dominance of incumbent hegemons (and their ethical priorities) as an unchangeable premise, seeking to ameliorate the prevailing system of power rather than challenging its foundations. 

Therefore, a successful Data Realist state must foster a "permanent view of politics" required to integrate the trajectory of global technological developments into its own strategic calculus, prioritize the development of sovereign technical and even management standards surrounding data, and explicitly link digital industrial geographies to national security goals. Just as the world has winners and losers, the digital society has data powers and data provinces.

3) The world leaks through.

Data realism is not a defense of surveillance, dashboards, spreadsheets, or technocratic governance. It is a defense of reality as something external to human discourse and design — something that can resist, surprise, and falsify our models. To that end, Data Realism rejects two dominant trends:

   Naive Empiricism — the idea that data “speak for themselves,” that numbers are neutral, measurement is innocent. This view fails to account for context, biases, or interpretation.

   Radical Constructivism — the view that data are nothing but power-laden constructs, shaped entirely by ideology, narrative, and positionality. This view erases the world and collapses epistemology into politics, often for sake of it.

A realist stance rejects both the blind faith in datafication and the nihilism of pure relativism. Data realism does not deny context, ideology, or structure. It insists that, even through those, the world leaks through. Consider a temperature reading, a mortality rate, or a vote count. These are not just narratives. They constrain us. To treat data as real is to take them seriously — not as final truth, but as our provisional contacts with the world. It is to ask what all this shows and means, not just who made it and why. Data therefore must be analyzed without idealization, where a commitment to the hard facts of data, even if inconvenient or ugly (e.g., showing inequality, corruption) is necessary and statistical gaslighting civilisationally poisonous. Data can be manipulated. But manipulation presupposes a baseline that can be distorted. Falsifying a vote count still depends on the idea of a real vote count. Censoring mortality rates still implies that there were deaths. To lie with data is to admit that truth matters, because at some level data are a non-negotiable reality that exist and operate independent of our political beliefs and moral aspirations. This means we must apply the highest scrutiny to the data used to train and test our systems, human or artificial. 

4) Data drive agency in the world.

Data is not an end but an index of industrial, military, and academic capacity. A state’s true data capacity is measured not by the size of its population’s data footprint, but by its independent ability to standardize, store, and compute that data without reliance on external supply chains or governance frameworks. This requires a systemic integration of military, academic, and industrial objectives — a union of science with industry that treats digital technical standards as global public goods that must be wielded strategically, and not just consumed passively. Contrary to idealistic claims, national security is the ultimate policy engine driving data governance decisions at the level of states, with privacy and ethics serving as secondary and often negotiable constraints. The data policies of the hegemonic and emergent powers are fundamentally rooted in securing technological advantage. The task therefore is not to eliminate dependency through isolation, but to gain the necessary leverage in data systems to shape the rules of its game.

Idealistic data policies are politically naive because they assume consent and cooperation in an anarchic system. Consider the G7's DFFT narrative, for instance, which is often projected as a universal good, but is mostly an elegant rhetoric of "free flow with trust" that uses an abstract legal promise as a mask to hide the concrete realities of global political controls left unacknowledged. A Data Realist state, therefore, must subject all policies to a simple but rigorous test: Does this policy measurably increase sovereign capacity and reduce structural dependency, or does it merely achieve moral compliance with the powers that be? Data Realism thus demands a meaningful shift from the judicial-police state (focused on making and enforcing laws) to the structural state (focused on building and owning digital capability and metapolicy spaces). 

5) Navigating the world with data requires pragmatism. 

Data realism is commitment to practical ethics, not idealisms. It eschews notions of diagonally opposite left/right systems. It is a philosophy of effective technological acceleration and not of technological pessimism. Practical ethics require direct and immediate confrontation with ethical necessities of data flows - make the methods of collection, cleaning, modeling, and interpretation as transparent as possible to those who are affected by the resulting decisions - but beyond a right to audit and redress, data stewards should not have to project desires of how the world or social contracts ought to be into their data pipelines. Data ownership is the ownership of Truth, and thus carries a responsibility to protect and de-risk the data in their care, and if required, transfer that ownership for systemic continuity. The primary task of governance here is thus not to make data and systems ethical, but to confront and master the measurable structural facts of computation, ownership, and capability. 

As almost everyone knows, the world and its governments are secretly run by accountants. This implies that data should go through continuous assessment for depreciation or appreciation. Once you put a number on the decay or change in data's subjective value due to context shifts and bias development, it will better incentivise the appropriate and timely flow of organisational resources to update and responsibly address the state of its data pipelines - as necessary to hold on to the realism in data. The financial incentives for better data governance and lifecycle management, reducing long-term infrastructural and technical debts, should thus be made explicit and immediate to the accounting elites. 

To conclude, as AI and automated systems gain more and more leeway into human affairs, this manifesto calls for embracing Data Realism, a philosophy anchored in the undeniable existence of the world and data's fundamental role as our provisional contact with it. It is a mandate to recognise that in a geopolitically volatile world, reliance on external, proprietary, or geographically constrained data sources can be a major systemic vulnerability - and argues for a strategic shift toward data resilience, reliability and sovereignty to guarantee uninterrupted operational continuity of our digital lives regardless of external regulatory or political pressures.