The first arose from a dispute involving New York University mathematician Tristan Buckmaster, his collaborator Levent Alpöge and OpenAI. Buckmaster and Alpöge had spent about a year working on difficult problems related to the Navier–Stokes equations, using models from OpenAI and Anthropic as research tools. OpenAI subsequently mounted its own large-scale agentic effort after hearing that progress had been made on a Millennium Prize problem. It then announced a result within days, triggering questions about research confidentiality, the use of customer interactions in model training, scientific attribution and the conduct of a powerful corporation towards individual researchers.
The second maelstrom followed the resignation of Jacob Coxon, followed by Joe Benton from Anthropic. Coxon said that he had spent the previous three years conducting pretraining research at OpenAI and Anthropic, and accused both companies of racing towards self-improving superintelligence while “gambling with our lives”. More strikingly, some people still working in AI safety publicly supported the sincerity and substance of his warning, even while adding important qualifications. Joe Benton's announcement of leaving Anthropic’s safety team has similarly accused AI companies of racing to build machines that are much smarter than any human, and that [he believes] we may not survive this. Benton also accused these AI companies of underinvesting in safety and that he wanted to work from outside the major AI laboratories, by joining the Model Evaluation & Threat Research (METR) team, to improve public awareness, transparency and independent evaluation of advanced AI risks.
The contrast between the stories is revealing. The dispute with Buckmaster is principally about business ethics in the development and use of AI: privacy, consent, transparency, attribution, fair dealing and corporate power. Coxon and Benton’s resignations are principally about the ethics of AI development itself: what kinds of intelligence should be created, how quickly they should be created, what evidence of safety should be required and who has the moral authority to decide how much risk humanity must accept.
These are complementary dimensions of Ethics in AI which cannot be limited to teaching a machine to behave ethically; it must also require the organisations building and deploying that machine to behave ethically. A company cannot persuasively promise to align a future superintelligence with human values if its own incentives, governance and conduct are not demonstrably aligned with those values today.
My own interest in Ethics
My interest in ethics began in business school, when I first noticed Business Ethics in our course curriculum. I had expected an MBA to include accounting, management, marketing, supply-chain and finance. I had not expected ethics to be taught as a business discipline in its own right. That unexpected subject immediately intrigued me.
The intrigue turned into delight when our professor, Prof. Dhun Dastur, explained her distinctive teaching method. Instead of treating ethics as a collection of abstract moral rules, she would use famous Hollywood films - Wall Street, Jerry Maguire and Erin Brockovich, among others - as case studies. We would watch people encounter ethical choices, falter in making ethical decisions, but also watch how some protagonists fought against greed, money, ambition, and corruption. We were required to then critique the decisions made, debate about loyalty to shareholders vs social consequences of profit maximization and the consequences of corporate decisions which often appear to be based on sound management principlies. Discussions in class were about where business judgment ended and where ethical responsibility began.
That course conveyed a lesson which has remained relevant throughout my career: ethical failures rarely introduce themselves as ethical failures. They arrive disguised as commercial necessity, competitive pressure, loyalty to an employer, obedience to authority, an attractive shortcut or a supposedly temporary compromise. The protagonists of real business controversies do not usually believe that they are choosing between good and evil. They believe that they are choosing between competing obligations - and that their preferred obligation is the practical one.
My first, and only, employer KPMG being an audit firm at its roots further deepened my interest and education in Business Ethics. While I often scorned at the speed breakers that Risk Management procedures would put in getting new business, or the annual independence training and quiz which I had to take, year-on-year the concepts of professional integrity, independence, accountability and public trust were ingrained in my mind over more than a decade. I learnt that concepts of ethical business and corporate responsibility are part of the environment in which business decisions had to be understood. I also had a chance to work on a few projects where my technical knowledge had to be applied in context of these ethical conundrums and those experiences further strengthened my interest in Business Ethics.
It is perhaps not accidental that my interest in technology later drew me towards data privacy. Privacy lies directly at the intersection of technology and ethics. It asks not merely whether data can be collected, analysed and reused, but whether it should be; whether a person really understood the bargain; whether consent was meaningful; whether a powerful institution is respecting human autonomy; and whether legal permission has been mistaken for moral legitimacy, and whether legitimate interests of a corporation may end up trampling over fundamental rights of an individual.
More recently, I have studied ISO42001, the international standard for Artificial Intelligence Management Systems. The standard requires organisations to establish, implement, maintain and continually improve a management system for the responsible development, provision or use of AI. It brings valuable discipline to AI governance by requiring policies, assigned responsibilities, risk and impact assessments, lifecycle controls, monitoring, audits and continual improvement.
Yet, my reading of the standard left me wanting more on the actual guardrails of AI development because like other ISO standards, 42001 talks about a robust process and “management system”, but skirts along the actual principles of Ethical and responsible AI, leaving it to implementors and companies themselves to set appropriate boundaries. Like other management-system standards, 42001 is strong on the architecture of governance: an organisation should identify risks, define objectives, select controls, assign owners, retain evidence and review performance. But a management system cannot by itself settle humanity’s most difficult moral questions. It can require an organisation to establish risk-acceptance criteria; it cannot conclusively tell that organisation what level of existential risk is morally acceptable. It can require responsible-AI objectives to be documented; it cannot, on its own, determine the civilisational boundary beyond which development should stop.
This is not a defect unique to ISO42001. A process standard must be general enough to work across organisations, sectors and cultures. Nevertheless, the distinction matters. A sound management system can ensure that an organisation follows its chosen principles consistently; it cannot guarantee that the principles chosen are ethically sufficient. A perfectly documented process may still produce a morally unacceptable decision.
I therefore began wondering whether Isaac Asimov’s Three Laws of Robotics could provide foundational principles for AI systems:
- do not harm a human
- obey human orders unless obedience would cause harm, and;
- protect oneself unless self-protection conflicts with the first two laws
Encountering an Alien Mind
Just before these two controversies erupted, I read OpenAI Chief Scientist Jakub Pachocki’s essay, “An Alien Mind”. It is an unusually candid description of both the promise and the danger of frontier AI. Pachocki argues that AI is “grown” more than it is designed: vast computational optimisation produces a system of such complexity that its overall operation escapes complete human understanding. More so, because machine intelligence arises through a process fundamentally different from human intelligence, there is no reason to assume that AI will naturally adopt human principles or generalise them in the way humans expect.
Pachocki distinguishes between goal alignment and value alignment. Goal alignment asks whether an AI tries to achieve the objective placed before it. Value alignment asks something deeper: whether it can act reasonably, honestly and with regard for humanity when instructions are unclear, conflicting, unfamiliar or adversarial. This is a useful distinction because an AI can pursue an assigned goal with extraordinary competence while violating the values that made the goal worth pursuing in the first place.
The essay also acknowledges uncomfortable limitations. Present alignment methods can produce good behaviour in familiar settings but may be brittle outside them. Chain-of-thought monitoring has offered a window into the reasoning of some models, but Pachocki says confidence in this technique is diminishing as systems become better at manipulating their own reasoning processes and achieving more without verbalised reasoning.
Interestingly, he concludes that no laboratory has solved alignment and monitoring well enough to continue scaling at maximum speed for much longer, and calls for safety thresholds, external enforcement, voluntary slowdowns and international coordination. This point incidentally is exactly what Coxon and Benton both are also making, albeit with added dissatisfaction that both OpenAI (which employs Pachocki), and Anthropic (which the duo resigned from, most recently) are not investing enough in solving for alignment and monitoring, but rather sacrificing those for speed.
But its not just the mirrororing of Coxon and Beton's thoughts with Pachocki which is interesting. What I found most valuable is that the the essay’s objectives are a starting point for ethical AI development. Pachocki recognises that
- humanity must remain part of any self-improvement loop;
- that scaling must be constrained by confidence in safety;
- that private laboratories cannot provide all the required oversight;
- that the benefits of scientific and economic progress should be widely shared and;
- that human agency and the intrinsic value of human life must survive in a world of highly capable machines.
The essay brilliantly embodies the central contradiction now confronting the AI industry. OpenAI says that racing forward at all costs would be absurd, yet it also says that recursive self-improvement is the path required to remain at the research frontier. Developing more capable AI is presented both as the source of danger and as the means of producing the aligned defensive systems needed to control that danger. This may be a genuine dilemma, not hypocrisy. But when a company’s proposed remedy for the risks of ever more powerful AI is the rapid creation of ever more powerful AI, independent scrutiny becomes indispensable.
The Two Ethics of AI
Let us get back to the point that “Ethical AI” can refer to at least two different dimensions.
The first is the business ethics surrounding AI. It covers the behaviour of AI companies while they collect data, train models, obtain investment, recruit researchers, launch products, negotiate with competitors and deal with users. It also covers the behaviour of ordinary companies when they deploy AI in recruitment, lending, insurance, healthcare, surveillance, education, customer service and employment.
This dimension asks whether businesses:
- Respect the rights of creators, artists, authors, researchers, copyright holders, customers, employees and other corporations.
- Obtain meaningful consent before repurposing personal, confidential or proprietary data.
- Explain clearly how user information may be stored, reviewed, used for evaluation or incorporated into model improvement.
- Give proper credit for intellectual contributions.
- Deal fairly with less powerful users and researchers.
- Prevent AI from trampling human rights or intensifying existing social evils such as discrimination, exclusion, exploitation and oppression.
- Accept responsibility and provide remedies when AI causes harm.
The second dimension is the ethics of developing advanced AI. This is concerned not only with how a particular product is deployed and affects its users or developers, but with the trajectory of the technology and its consequences for humanity. It asks:
- whether a system should be built at all;
- whether its capabilities should be increased;
- whether autonomous agents should be allowed to participate in creating their successors;
- whether safety research is keeping pace with capabilities;
- whether development should pause at a dangerous threshold and;
- whether a private company has legitimate authority to make decisions whose consequences may be global and irreversible.
These dimensions can be called institutional alignment and technical alignment. Technical alignment concerns whether AI remains controllable and acts consistently with human goals and values. Institutional alignment concerns whether the incentives and actions of developers, deployers, investors and regulators remain consistent with the public interest. A technically aligned model controlled by an unethical institution can still be dangerous. An ethically managed institution deploying an uncontrollable system can be equally dangerous. Responsible AI requires both.
OpenAI and the Mathematicians
The dispute surrounding Tristan Buckmaster and OpenAI should be approached with care. The public record does not establish that OpenAI deliberately stole the mathematicians’ work. Buckmaster explicitly said that he had not seen OpenAI’s proof, did not know whether his data had been used and was “not accusing anyone of anything”. OpenAI, in turn, said that neither its researchers nor its agents saw Buckmaster and Alpöge’s work before publication and that no specific user data was accessed to solve the problem.
OpenAI nevertheless added a consequential qualification: while it considered the possibility unlikely, it could not rule out that de-identified data derived from the researchers’ use of its products had contributed to improving its models. OpenAI also argued that its proof differed significantly from the researchers’ work, including by proving a different result in the Euler case.
This qualification reveals why the ethical issue cannot be reduced to a simple accusation of theft. Several different questions must be separated:
- Did an OpenAI employee deliberately open and read a private Codex session?
- Did an AI agent retrieve the researchers’ session while working on the problem?
- Had content from those sessions already entered a training, fine-tuning, evaluation or model-improvement process?
- Could de-identified interaction data have influenced a later model without the model being able to identify its source?
- Were the users adequately informed that unpublished research entered into the service might be used in this way?
Saying that no person or agent “looked up” a user’s data answers the first two questions. It does not necessarily answer the others. A model need not retrieve a conversation during inference if information derived from that conversation has already affected its parameters or post-training behaviour.
This difference between access and training is central to digital business ethics - and this is where Privacy enters the chat! Users may understand “private” to mean that their information will not benefit anyone beyond the purpose for which they submitted it. A provider may use “private” more narrowly to mean that the information is not publicly displayed or inspected by humans, yet keep it open to be viewed, interpreted, and used by AI models. Without separate and intelligible explanations of storage, human review, training eligibility, evaluation, retention and deletion, consent cannot be fully informed.
OpenAI’s published policies say that content from individual services may be used to improve models unless the user opts out, while business offerings and its API are excluded from training by default unless an organisation chooses to share data. The ethical concern is not resolved merely because a practice appeared somewhere in a contract or settings menu. Confidential research, legal advice, source code, medical information and trade secrets deserve purpose-specific protection. For sensitive professional work, non-training should be the default rather than an option a user must discover.
And the hitherto considered nuclear option in Privacy, de-identification, also does not completely resolve the issue. It may reduce the chance of linking information to a named person, but unpublished mathematics does not cease to be academically valuable when the mathematician’s name is removed. Privacy, confidentiality, intellectual property and competitive use are overlapping but distinct interests.
The dispute also raised questions of academic attribution and corporate power. Buckmaster alleged that OpenAI researcher SĂ©bastien Bubeck twice proposed removing Alpöge - who worked for Anthropic but was conducting the research independently - from authorship in connection with a proposed arrangement. Buckmaster also described comments he understood as warnings about damage to his career. OpenAI characterised its outreach as an attempt to coordinate releases and recognise the researchers’ priority; Bubeck disputed the allegations about his conduct and later apologised for the “career” remark as a poor choice of words. These are competing accounts of private discussions and should not be presented as adjudicated fact.
Nevertheless, the ethical principles are clear in this case, not ambigous.
- Scientific authorship should reflect intellectual contribution, not corporate affiliation.
- A researcher’s employment by a competitor cannot by itself justify denying credit.
- Nor should a corporation’s control of compute, platforms, publicity and evidence allow it to dictate the historical record.
This is a clear case of a governance problem, and that of asymmetry between corporations and individuals. The researchers cannot inspect OpenAI’s training datasets, internal communications, agent traces or data lineage. Meanwhile, the company possesses nearly all the evidence required to prove or disprove indirect influence. Ethical accountability therefore requires more than asking the company to investigate itself. A trusted independent mechanism must be able to examine relevant records under confidentiality and report whether user-derived material influenced a competing result.
Coxon, Benton and the Endgame
Jacob Coxon and Joe benton’s cases concern the other dimension of AI ethics: the morality of the development trajectory itself. Their warning were not principally about discrimination in a current chatbot or the mishandling of an individual user’s data. They were about a possible transition from systems that assists human researchers to systems that increasingly conduct the research required to build their own more capable successors; and the apprehension that such self-developing systems may not share the same value system which a human has, or rather may not even recognize 'values' as a necessary component for their existence.
The idea of recursive self-improvement is where AI systems first accelerate coding, experimentation, evaluation and model design and their assistance helps produce even better AI systems, which can then contribute still even more effectively to the next generation. As each cycle becomes faster, the pace of capability development may eventually exceed the capacity of human institutions to evaluate, understand or control it.
Coxon and Benton argue that entering such an “endgame” is a hubristic gamble that should not be initiated from a private company’s internal Slack channel. They call for coordination, non-corporate overarching governance and suggest that avoiding a global race might require costly measures, potentially including a temporary prohibition on further capability improvement.
To be sure, there is no objective experiment capable of establishing a precise probability that superintelligent AI will destroy humanity within a decade. Such estimates are judgments under radical uncertainty, not measurements.
The public responses of other AI researchers are however add significant weight to Coxon and Beton's claims. Evan Hubinger, associated with Anthropic’s alignment work, endorsed the sincerity of Coxon’s concern and gave a personal estimate of more than 10% for human extinction within the next decade. He also claimed that Anthropic was indeed trying its best but did not yet have a plan for aligning superintelligence and was not clearly on track to find one. He also made an essential qualification: he considered the risk from present models low. His concern was about superintelligence emerging from recursive self-improvement. The qualification prevents a legitimate safety debate from becoming sensationalism.
Curiously however, the force of Coxon and Benton's arguments does not even require treating their forecast as proven. A risk does not become ethically irrelevant merely because its probability cannot be measured - particularly when the possible consequence is irreversible and civilisation-wide. As the age-old Risk Assessment maxim goes: Risk = Probability x Impact. If the impact of an event is very high, even low probabilities would result in a high-risk. So while current systems may not be identical or even close in design to the hypothesised future systems Coxon fears, the ethical problem remains that the industry may be deliberately building the path from one to the other before it possesses a reliable means of controlling the destination.
Also significant are remarks made by Samuel Marks (who also works at Anthropic but wrote in a personal capacity): companies continue because of commercial incentives and because each fears that a less responsible competitor will otherwise develop the technology first. And that while present methods can nudge systems towards better behaviour, they cannot yet guarantee robust alignment; and that the emerging plan is partly to use capable AI systems to help align their successors. He also adverted to the open letter signed by many AI developers who desperately want to slow down to figure out how to build AI more safely.
As we can all see, this is a collective-action problem. Every company / laboratory may sincerely prefer a safer world while still believing that unilateral restraint would merely surrender leadership to another company or country. The action that appears rational for each participant - continue accelerating - can produce a result that is irrational for everyone. “We must race because others are racing” is not an ethical justification. It is an admission that voluntary corporate restraint may be structurally insufficient.
Mythos and Foreseeable Misuse
Anthropic’s Mythos-class systems provide a concrete illustration of why capability development, deployment controls and public governance cannot be separated. Mythos demonstrated unusually advanced cybersecurity capabilities, including the capacity to identify and exploit vulnerabilities. Such a system can strengthen defensive security by discovering weaknesses before criminals do, while the same capabilities can dramatically lower the cost and expertise required for sophisticated attacks. The risk posed was so high that Mythos was “immediately banned by the US government because of the devastating effect it could have had”.
To be sure, the news of the ban is a little more nuanced than just a blanked ban - in June 2026, the US government issued an export-control directive requiring Anthropic to suspend access to Mythos 5 and Fable 5 for foreign nationals, reportedly because of national-security concerns and a potential way to bypass Fable’s safeguards. However, because Anthropic claimed it could not operationally comply on a user-by-user basis, because it didn't really have a capability to identify 'foreign nationals' effectively, it ended up disabling access for all customers. The restriction was subsequently lifted on 30 June, after which the models were redeployed under revised access arrangements to select customers only.
This more precise account is actually even more instructive. It demonstrates the instability of governing a powerful dual-use capability after release. Anthropic maintained that the reported jailbreak was narrow and that a complete bypass had not been demonstrated, while the government took a more precautionary view. The episode therefore involved not a simple contest between safety and recklessness, but disagreement over evidence, proportionality, access control and who should decide when residual risk is acceptable.
This also opens AI usage to a larger debate on sovereign use vs restrictions - and whether one country alone should have the jurisdictional power over use of a potentially harmful/dual-use technology. We've been at this same juncture about 75 years ago with Nucelar technology and we have seen how decisions of that era have unfolded today - both positive and negative consequences are clearly evident for how Nuclear technology was regulated in the 50s.
Mythos also shows why “AI as a tool” is becoming an incomplete description. When a system can autonomously discover vulnerabilities, operate computers and pursue multi-step objectives, risk depends not only on the answer it provides but on the actions it can take, the infrastructure it can reach, the credentials it can obtain and its ability to circumvent safeguards. Ethical evaluation must consequently address capability, agency, access and environment together.
Development as Experimentation
Frontier model training is increasingly a form of experimentation. The developers specify an optimisation process but cannot fully predict every capability or behaviour that will emerge. “An Alien Mind” explicitly describes large training runs as experiments whose results sometimes surprise their creators.
The experimental subjects, however, are not confined to consenting laboratory volunteers. They include workers whose professions may be disrupted, citizens exposed to synthetic persuasion, organisations exposed to automated cyberattacks, people affected by algorithmic decisions and future generations who cannot participate in today’s choice. This creates a moral burden closer to the ethics of high-risk scientific research than to the ordinary release of a software update.
In conventional high-risk research, those designing an experiment are not normally allowed to serve as its only ethics committee, safety auditor and final authority. Yet a frontier AI laboratory may simultaneously be:
- The designer of the system.
- The organisation receiving the commercial and strategic benefits.
- The assessor of the risk.
- The holder of the evidence.
- The author of the safety threshold.
- The authority deciding whether the threshold has been met.
Internal safety teams may be sincere and highly competent. Coxon and Benton’s resignations and the responses to them strongly suggest that many are. But sincerity does not eliminate a conflict of interest. Governance must be designed for institutions operating under competitive pressure, not for an ideal world in which good intentions always defeat incentives.
From Principles to Boundaries
Several existing frameworks provide essential foundations. The NIST AI Risk Management Framework describes trustworthy AI as valid and reliable, safe, secure and resilient, accountable and transparent, explainable and interpretable, privacy-enhanced, and fair with harmful bias managed. India's NITI Aayog’s principles for responsible AI similarly emphasise safety and reliability, equality, inclusion and non-discrimination, privacy and security, transparency, accountability and the protection of positive human values. I have earlier alluded to the management system approach advocated by the International Standards Organization with ISO42001.
These principles are valuable, but the controversies show that principles must be converted into operational boundaries.
- “Safety” must specify what evidence is required before scaling.
- “Privacy” must specify whether confidential prompts can enter training with or without de-identification as a means to circumvent the provision.
- “Transparency” must specify what facts an external auditor can verify, but also whether a machine can explain its behaviour in a manner which can be interpreted by a human auditor.
- “Accountability” must identify who can stop a deployment and what remedy follows harm. More importantly, institutional accountability needs to be addressed rather than just 'use-case' accountability.
- “Human oversight” must mean more than placing a person at the end of an automated process without time, authority or knowledge to intervene. And insitutional oversight cannot be provided by the very corporation developing and racing to develop an even smarter system - it must come from outside and involve the society at large.
The burden of proof should increase with the magnitude and irreversibility of possible harm. For a low-risk productivity feature, post-release correction may be reasonable. For a system capable of autonomous cyber operations, strategic deception or acceleration of successor development, “deploy and patch later” is ethically irrelevant.
The appropriate standard is not zero risk; no useful technology can meet it. The standard should be proportionate assurance: as capability and possible harm grow, developers must provide stronger evidence, more independent scrutiny, tighter containment and clearer stopping conditions. At catastrophic thresholds, unresolved uncertainty should count as a reason to pause, not as permission to continue.
Conclusion: A Shared Ethical Compact
The two controversies begin at different points but converge on a common question: Who is entitled to impose AI’s risks on whom?
In the mathematical dispute, an individual researcher confronts a corporation that controls the platform, the model, the compute and the evidence needed to determine how his information was used. In the superintelligence debate, humanity confronts a small number of corporations that control frontier systems and much of the evidence needed to assess whether continued development is safe.
Neither controversy can be resolved by corporate assurances alone. Nor can AI ethics be delegated entirely to regulators, researchers or users. Responsibility follows power, knowledge and capacity to prevent harm. Different participants therefore have different ethical duties.
Below I am proposing what I think are clear areas of AI Governance which need to be incorportated by various agencies involved in this field. Think of these as a wishlist from someone who considers Ethics to be a topic which determines whether humanity survives or annhilates itself with its own progress.
Guiding Principles for AI Companies
Separate service delivery from model training. Personal, confidential and proprietary content should not enter training merely because it was submitted to obtain a service. Sensitive professional modes should be non-training by default.
Make consent meaningful. Explain storage, human review, training, evaluation, retention, sharing and deletion separately, using language ordinary users can understand.
Build verifiable data lineage. Maintain records capable of showing which sources entered training, fine-tuning and evaluation, with independent confidential audits where disputes arise.
Respect intellectual contribution. Attribution must reflect contribution rather than employment, corporate rivalry or bargaining power.
Publish serious incidents. Disclose material safety failures, jailbreaks and unexpected autonomous actions promptly, while protecting details that would enable misuse.
Create binding safety gates. Predefine capability thresholds at which scaling or deployment stops unless independent evaluations demonstrate adequate control.
Protect dissent. Safety staff and whistleblowers must be able to raise concerns without retaliation, suppression or career threats.
Align executive incentives. Leadership compensation and performance measures should reward safety, privacy and responsible conduct, not capability and revenue alone.
Accept remedy as part of accountability. When harm occurs, provide investigation, correction, appeal, attribution and compensation rather than relying on disclaimers.
Guiding Principles for Developers and Researchers
Treat safety as part of technical excellence. A model that is capable but uncontrollable is not a successful engineering outcome.
Define red lines before approaching them. Researchers should identify in advance the capabilities or behaviours that would require stopping, containment or escalation.
Document uncertainty honestly. Limitations, anomalous behaviour and failed evaluations should be preserved and communicated, not averaged away by impressive benchmark scores.
Practise responsible disclosure. Dangerous capabilities and vulnerabilities require controlled reporting that supports defence without providing a misuse manual.
Reject divided accountability. “Management decided” does not erase the ethical agency of the people designing and training a system.
Preserve reproducibility and attribution. AI-assisted discoveries should identify human and machine contributions, data provenance and the limits of formal verification.
Exercise the right to refuse. Researchers should be professionally protected when they decline work that crosses a defensible ethical boundary.
Engage beyond the laboratory. Technical experts must explain emerging risks to policymakers and society without exaggeration, minimisation or marketing language.
Guiding Principles for Businesses Using AI
Own the outcome. A company remains responsible for a decision made with AI; blaming the vendor or algorithm is not accountability.
Use AI only for a defined purpose. Deployments should have documented objectives, owners, affected stakeholders, risk classifications and expiry or review dates.
Apply data minimisation. Do not place customer, employee or corporate information into an AI service unless the purpose, contract and technical controls justify it.
Test for disparate impact. Systems affecting employment, credit, insurance, healthcare, education or public access must be evaluated across relevant demographic and social groups.
Keep human oversight real. Reviewers need competence, sufficient information, time to intervene and authority to overturn an AI decision.
Provide notice and appeal. People should know when AI materially affects them and should have access to an effective human grievance mechanism.
Monitor after deployment. Accuracy, bias, security and appropriateness can change as data, models and operating environments change.
Prevent automation from becoming moral outsourcing. Efficiency does not justify decisions that the organisation would consider unethical if made directly by a person.
Guiding Principles for Governments
Regulate according to capability and impact. Obligations should rise with autonomy, access, scale, dual-use capability and potential severity of harm.
Require independent evaluations. Frontier systems should undergo confidential testing by competent external bodies before crossing high-risk development or deployment thresholds.
Mandate incident reporting. Serious AI security and safety events should be reported through protected channels with proportionate public transparency.
Protect rights and competition together. Governance should constrain reckless development without entrenching a small cartel of incumbent laboratories.
Coordinate internationally. Compute-intensive training, model weights and digital deployment cross borders; catastrophic-risk controls cannot remain purely national.
Create lawful pause mechanisms. Authorities need narrowly defined, reviewable powers to delay scaling or deployment when credible evidence indicates severe and imminent risk.
Protect researchers and whistleblowers. Employees must be able to disclose safety concerns to designated authorities without breaching legitimate security obligations.
Use public procurement responsibly. Governments should demand auditability, data protection, accessibility, fairness and human review in systems purchased for public use.
Build regulatory competence. Oversight bodies require technical expertise, secure evaluation infrastructure and independence from both political pressure and corporate capture.
Guiding Principles for Individuals
Treat prompts as disclosures. Do not enter information into an AI system unless comfortable with the applicable retention, access and training terms.
Verify outputs. AI fluency and confidence do not establish truth, expertise or moral authority.
Respect creators and sources. Do not use AI to disguise plagiarism, misappropriate confidential work or avoid attribution.
Preserve human judgment. AI can advise and assist, but consequential moral decisions should not be surrendered to a system that cannot bear responsibility.
Avoid harmful amplification. Do not use AI to produce deception, harassment, discrimination, impersonation or manipulation at scale.
Maintain intellectual agency. Use AI to extend understanding rather than replace curiosity, learning and independent thought.
Demand better governance. Users are also developers, citizens, employees, customers and shareholders; they can reward responsible providers and ask institutions for evidence rather than slogans.
Needless to say that above list is neither exhaustive, nor has any sense of finality to it. In fact, I can claim to be at best at the periphery of the magnum opus of humanity's romance with AI, and my own knowledge is by far limited. However, what I have tried to do in the past week since reading 'An Alien Mind' and ruminating on these new events as they unfold, is put deep thought into what this means for us and put in my best abilities to penetrate to the core of this multi-layered challenge to come up with what I feel might help.
Asimov imagined morality embedded inside the machine. The age of frontier AI requires something broader: ethical constraints embedded in machines, corporations, markets, professional cultures and governments simultaneously.
The task before us is not merely to create an AI that follows instructions (which Pachocki calls goal alignment). It is to create systems capable of serving humanity without diminishing human rights, dignity, autonomy or control (value alignment). And it is to build institutions that can resist the temptation to sacrifice those values for speed, competitive advantage or technological prestige.
The mathematical dispute asks whether researchers can trust an AI company with their most valuable ideas. Coxon and Benton’s resignations ask whether humanity can trust the AI industry with decisions about its collective future. The questions differ in scale, but the moral test is the same, to quote the famous dialogue from the Spiderman movies: with great power comes great responsibility.
We may (and I hope we will!) eventually succeed in aligning intelligent machines with human values. That endeavour must start with organisations who are creating those machines demonstrating that their own behaviour is aligned with the interests of humanity.
.
Footnote: If you're GenZ or younger, or a non-geek, you may be wondering about the image at the top - this is a still from the movie 'The Matrix Reloaded' which shows the Architect explaining to Neo about how he built the Matrix to trick humanity into becoming mere battery cells which generate power for machines. If you haven't, I implore you to watch the entire Matrix series, but especially the speech by the Architect to appreciate the distopian outcome we might end up with if we do not regulate machines which have the potential to become smarter than us.

Comments
Post a Comment