Showing posts with label Artificial Intelligence. Show all posts
Showing posts with label Artificial Intelligence. Show all posts

Saturday, September 26, 2026

SOMEBODY STOP US: Who Will Control Artificial Intelligence?

SOMEBODY STOP US: Who Will Control Artificial Intelligence?

Is artificial intelligence really slipping out of control, or are we moving toward a different kind of danger: a future in which only a handful of governments and technology giants are able to develop the most advanced AI systems?

For me, the central question is no longer simply, “How dangerous could artificial intelligence become?”

The more important question is:
“Who will control artificial intelligence, and could that control mechanism eventually turn into a monopoly that decides who is allowed to develop AI?”

2023: The First Serious Alarm Bells

Today’s debate is not entirely new. The year 2023 marked an important turning point in discussions about the future of artificial intelligence.

Geoffrey Hinton, one of the pioneers of modern AI, left Google in 2023 and began speaking more openly about the risks of artificial intelligence. He warned that AI systems were developing much faster than he had expected and that increasingly capable systems could create serious risks in the future.

In the same year, an open letter published by the Future of Life Institute made an even more extraordinary proposal.

Among its signatories were Elon Musk, Yoshua Bengio, Steve Wozniak and many researchers and technology leaders.

The letter called for at least a six-month pause in the training of AI systems more powerful than GPT-4.

At the time, many people considered the proposal unnecessarily dramatic.

The dominant logic of the technology industry was exactly the opposite: build larger models, use more computing power, collect more data, and push performance further.

Three years later, however, the picture looks quite different.

2026: This Time, the AI Companies Themselves Are Sounding the Alarm

Today, concerns about AI safety are no longer being raised only by academics, activists or technology critics.

Some of the companies leading the AI race are themselves calling for stronger safety and governance mechanisms.

OpenAI has supported mandatory, capability-based national AI safety rules. Anthropic has argued that a technology developing as rapidly and affecting society as broadly as artificial intelligence should not be governed by industry alone, and that governments also need enforceable rules. Google DeepMind, through its Frontier Safety Framework, has established critical capability thresholds for areas including cybersecurity, biological risk, autonomy, manipulation and loss of control.

In other words, the debate is no longer purely theoretical.

The companies developing the most capable AI systems are themselves evaluating scenarios in which future models could enable powerful cyberattacks, biological threats, large-scale manipulation or serious control problems.

And the concern is no longer limited to malicious humans using AI.

Increasing attention is also being paid to highly autonomous systems that may pursue unintended objectives, circumvent human supervision or behave in unexpected ways.

Therefore, it would be a mistake to dismiss the entire AI safety debate as manufactured fear. There are real technical and societal risks that deserve serious attention.

But Another Race Is Taking Place at the Same Time

This is where the second part of the problem begins.

Artificial intelligence is no longer merely a scientific or commercial technology.

It has become a strategic technology with implications for economic power, cybersecurity, defense, intelligence, scientific research and geopolitical influence.

This leads to another powerful argument, particularly in the United States:

“If we slow down, there is no guarantee that our competitors will slow down with us.”

The fact that the U.S. AI Action Plan prominently frames the issue in terms of “Winning the AI Race” illustrates how central this logic has become.

The plan emphasizes accelerating innovation, expanding AI infrastructure and maintaining American leadership in artificial intelligence.

Particularly in the context of competition with China, AI policy is increasingly beginning to resemble national security policy rather than conventional innovation policy.

This creates a remarkably difficult equation:

On one side:
More capable AI models could create serious risks. Testing, oversight, safety standards and, in some cases, restrictions may be necessary.
On the other side:
If one country or company slows down, there is no guarantee that its competitors will do the same.

My concern begins precisely where these two arguments intersect.

Could Safety Regulation Become a Barrier to Entry?

I also believe that artificial intelligence needs meaningful control mechanisms.

But I strongly disagree with the assumption that “AI needs control” must automatically lead to the conclusion that control should rest in the hands of a few large governments and a few major corporations.

Regulation creates a very important structural problem:

A safety requirement that is relatively easy for a trillion-dollar company to satisfy can become an impossible barrier for a startup, university laboratory or independent research group.

Imagine that developing a frontier model eventually requires extremely expensive licenses, mandatory certification, multimillion-dollar safety infrastructure, audits that can only be conducted by a small number of approved organizations, or special government authorization.

On paper, the rules might apply equally to everyone.

In practice, however, only a very small number of organizations might still be able to develop advanced AI.

Would such regulation reduce competition?
Would it make it harder for new companies to emerge?
Would it restrict frontier AI research at universities?
Would it ultimately strengthen the position of the companies already dominating the field?

I believe these questions deserve much more attention.

Why the Nuclear Technology Analogy Keeps Appearing

Artificial intelligence is frequently compared with nuclear technology.

The analogy has obvious limitations.

An AI model is not a nuclear reactor. Software can be copied, modified and deployed in many different parts of the world.

Yet when we look specifically at frontier AI, one important similarity is becoming increasingly difficult to ignore.

Developing the most advanced systems requires rapidly increasing amounts of computing power, energy, data-center infrastructure, advanced semiconductors, highly specialized human capital and enormous financial resources.

If heavy regulatory requirements are added on top of these existing barriers, frontier AI development may naturally become concentrated in the hands of fewer and fewer actors.

And that creates an uncomfortable relationship between safety regulation and the centralization of technological power.

Why Open-Weight AI Is at the Center of This Debate

One of the most important battlegrounds in this discussion will be open-source AI and, more precisely, open-weight models whose model weights are available for others to access and use.

Open models offer important advantages.

Universities can conduct independent research. Startups can develop products without becoming completely dependent on large API providers. Organizations can run models on their own infrastructure without sending sensitive data to third parties. Countries can develop systems tailored to their own languages, cultures and strategic needs.

At the same time, openness can also make it easier to remove safety mechanisms or modify highly capable models for malicious purposes.

Therefore, there is no simple answer to the question of whether AI should be open or closed.

But one distinction is essential.

Discussing the risks of open models is one thing.

Using those risks as a justification for creating an AI ecosystem in which everyone becomes dependent on a handful of closed platforms is something entirely different.

But There Is Also Important Counter-Evidence

If we want to evaluate this issue seriously, we should not only look for evidence that supports our concerns.

We should also look for evidence that challenges them.

At present, there is no clear evidence that the major AI companies are collectively pursuing a strategy to eliminate open AI.

Anthropic CEO Dario Amodei has stated that his company does not advocate banning open-weight models and has described open models without dangerous capabilities as a public good.

More importantly, the current U.S. AI Action Plan explicitly includes policies aimed at supporting open-source and open-weight AI.

The plan recognizes the value of open models for startups, academic research, organizations working with sensitive data, and the broader international influence of the U.S. technology ecosystem.

Therefore, it would not be justified to claim that there is already a proven plan to shut down open-source AI.

My concern is different.

We might reach the same outcome even without any secret plan.

Poorly designed incentives, excessively expensive safety obligations, concentration of computing power and national-security concerns could gradually create a system in which only a small number of actors are capable of developing frontier AI — even if nobody originally intended to create such a monopoly.

The Next Stage: AI Sovereignty

For me, this issue extends far beyond competition between companies.

The larger question concerns countries.

In the future, having strong AI researchers may not be enough to develop the most capable AI systems.

Countries may also need access to advanced GPUs, massive data centers, energy infrastructure, leading-edge semiconductor technologies, significant capital and international supply chains.

If international licensing systems, export controls or authorization mechanisms for frontier AI are added to this structure, technological sovereignty will become an even more important issue.

At that point, we may have to ask:

If a country needs another country's permission to develop one of the most important general-purpose technologies of the future, can it really be considered technologically independent?

This question will become increasingly strategic for every country that wants to be not merely a user of artificial intelligence, but also a producer of it — including Türkiye.

So Is the Solution “No Regulation at All”?

No.

I believe both extremes in the AI debate are problematic.

The first extreme is to allow technology to advance without meaningful oversight and simply expect the market to solve serious safety problems.

The second extreme is to use safety concerns to create a system in which the ability to develop advanced AI is effectively controlled by a handful of companies and governments.

What we need is a governance model somewhere outside these two extremes.

Regulation should be based on measurable capabilities and risks rather than company names or model size alone.

Audit criteria should be transparent.

Universities and independent researchers should retain the ability to conduct meaningful AI safety research.

Compliance costs should not automatically push startups and smaller research organizations out of the field.

Not every open model should be treated as belonging to the same risk category.

A small language model and a frontier system capable of advanced biological design or sophisticated cyber operations should not necessarily be governed by identical rules.

And perhaps most importantly:

The rules governing the future of artificial intelligence should not be written only by the AI companies themselves.

When the companies being regulated also become the primary architects of the regulation, a natural conflict of interest may emerge.

The Real Issue May Be Bigger Than Safety

The developments of the past three years have made one thing increasingly clear to me.

AI safety is a real issue and deserves serious attention.

But from now on, it will not be enough to ask whether AI models themselves are safe.

We also need to ask how economic, technological and political power over artificial intelligence is distributed.

Because an AI ecosystem that is safe but completely controlled by a very small number of actors may not be an ideal outcome for humanity either.

I believe one of the most important technology-policy debates of the coming years will therefore be:

How should we control artificial intelligence?

And more importantly:
Who will control those who control AI?

My concern, therefore, is not that we should avoid controlling artificial intelligence.

Quite the opposite.

We should build serious, scientific, transparent and democratically accountable mechanisms for governing advanced AI.

But while trying to make AI safer, we should be careful not to turn one of the most powerful technologies in history into the permanent privilege of a few governments and a few giant corporations.

Because one day, the biggest problem may not be that artificial intelligence has escaped our control.

It may be that control over artificial intelligence has become far too concentrated.


What do you think?
Do frontier AI systems require stronger global oversight? Or could the mechanisms created to ensure safety eventually produce an even greater risk by concentrating technological power in the hands of a few countries and corporations?

#ArtificialIntelligence #AISafety #AIGovernance #OpenSourceAI #OpenWeights #FrontierAI #AGI #TechPolicy #AIRegulation #AISovereignty #OpenAI #Anthropic #DeepMind #MuratKarakayaAkademi

Are Closed LLMs Losing Their Lead? Open-Weight Models Are Moving Closer to the Frontier

```

Are Closed LLMs Losing Their Lead? Open-Weight Models Are Moving Closer to the Frontier

Inkling, Kimi K3, and Qwen3.8 changed the open-weight LLM landscape within just a few weeks, raising important questions about the future of the AI race.

As we entered the summer of 2026, the strongest players in the large language model race still looked very familiar: OpenAI, Anthropic, and other closed-model providers. Anthropic released Claude Fable 5 on June 9, 2026, while OpenAI made the GPT-5.6 family generally available on July 9 following a limited preview that began on June 26.

Both models were introduced with strong reasoning, coding, and agentic capabilities. Their advanced capabilities and potential security risks, particularly in cybersecurity, also brought them into broader discussions involving the U.S. government.

Then, within only a few weeks, three major developments changed the picture on the open-weight side.

🗓️ A landscape that changed within weeks

🇺🇸 July 15, 2026 — Inkling: Thinking Machines Lab released Inkling, an open-weight model with 975 billion total parameters and 41 billion active parameters per token.

🇨🇳 July 16, 2026 — Kimi K3: Moonshot AI announced Kimi K3, a 2.8-trillion-parameter model. Its full model weights were released on July 27.

🇨🇳 July 19, 2026 — Qwen3.8 preview: Alibaba introduced Qwen3.8-Max-Preview in the 2.4-trillion-parameter class. Qwen3.8-Max was formally introduced on August 3, followed by the open-weight Qwen3.8-2.4T-A95B model on August 12.

🇺🇸 Inkling: A major open-weight move from the United States

The significance of Inkling is not simply that it has 975 billion parameters. The model uses a Mixture-of-Experts architecture with around 41 billion active parameters per token and can reason across text, images, and audio.

It also supports agentic coding, tool use, and context windows of up to one million tokens.

What makes Inkling particularly interesting, in my view, is something else. While Chinese companies had become far more visible in frontier open-weight model development, this time a powerful and genuinely open-weight model came from a new U.S.-based AI laboratory.

🇨🇳 Kimi K3: Open-weight models approach the 3-trillion-parameter class

Immediately after Inkling, Moonshot AI introduced Kimi K3, pushing open-weight model scale into a new range.

The model contains 2.8 trillion parameters. Its MoE architecture includes 896 experts, while only 16 are activated for each token. This allows the total model size to become extremely large without requiring every parameter to participate in every inference step.

Kimi K3 also attracted attention for more than its size. Around the time of its release, it ranked first on Frontend Code Arena, ahead of Claude Fable 5.

That does not mean Kimi K3 is now simply “better than GPT-5.6 or Fable 5.” Closed frontier models still hold an advantage in many broad evaluations. But the fact that an open-weight model can outperform the strongest closed models in selected coding and agentic tasks is a substantial change compared with the situation only a few years ago.

🇨🇳 Qwen3.8: Alibaba also moves into trillion-scale open models

Alibaba’s Qwen3.8 move showed that these developments were not isolated events.

Qwen3.8-Max, previewed in July, reached its formal 2.4-trillion-parameter release in August. The Qwen team then released the weights of Qwen3.8-2.4T-A95B as well.

As a result, within a short period, the open-weight ecosystem saw new models at 975 billion, 2.4 trillion, and 2.8 trillion parameters.

⚖️ The real issue is not parameter count

A 2.8-trillion-parameter model is not automatically better than a 500-billion-parameter model. Training data quality, architecture, post-training, reinforcement learning, tool use, inference infrastructure, and active parameter count can matter just as much as total model size. The real shift is that open-weight models are beginning to operate in the same capability range as frontier systems.

🔓 Are open-weight models still six or seven months behind?

For several years, there was a fairly reasonable rule of thumb: the strongest capabilities first appeared in closed systems from companies such as OpenAI, Anthropic, or Google, while comparable open models followed months later.

That assumption now needs to be reconsidered.

Models such as Kimi K3 can enter the same performance region as GPT-5.6 and Claude Fable 5 on selected tasks. The Qwen family continues to improve rapidly in coding and agentic workloads. Inkling shows that a U.S.-based open-weight frontier model is also possible.

Open models have not yet surpassed closed models across every dimension. But in many areas, the delay no longer looks like “wait for the next model generation.”

To me, this is the more important change.

🇨🇳 Are Chinese companies catching up with U.S. frontier labs?

Recent developments make this question increasingly difficult to dismiss.

Chinese companies are no longer producing only smaller or cheaper alternatives. With Moonshot AI and Alibaba, we are now seeing models at true frontier scale, and some of them are being released with their weights.

This is particularly significant because Chinese companies are operating under restrictions affecting access to some of the most advanced U.S. AI chips. Despite this, they continue to scale through work on model architecture, sparsity, attention mechanisms, quantization, and distributed inference.

The competition may therefore increasingly shift from “Which company has the best model?” to “Which national ecosystem can still develop frontier models?”

🔐 Will Chinese companies eventually close their strongest models?

This may become one of the most interesting questions for the future of the open-weight ecosystem.

Chinese model developers have benefited significantly from open-weight strategies. Their models have spread quickly across the world, developer communities have formed around them, inference platforms have added support, and the companies themselves have gained global visibility.

But would a company follow the same strategy if it genuinely believed it had the most capable model in the world?

Or would it become commercially more attractive to keep the strongest model behind an API while releasing the previous generation as open weights?

The next few model generations may provide an answer.

🌍 Where are the countries outside the U.S. and China?

This may be one of the least discussed aspects of the current AI race.

Two ecosystems now stand out very clearly in frontier model development: the United States and China.

Europe has strong research institutions and important AI companies. Japan and South Korea have advanced semiconductor industries. India has a very large software and engineering talent base. Yet at the trillion-parameter frontier, the competition is increasingly concentrated in the U.S. and China.

The reason is that this race cannot be run with good researchers alone. It requires tens of thousands of accelerators, data-center infrastructure, energy, capital, training data, distributed-systems engineering, and a very large research ecosystem.

One of the strategic questions for the coming years will therefore be whether other countries attempt to build their own frontier-model infrastructure or become increasingly dependent on models developed in the U.S. and China.

❓ Five questions for the next phase

🔹 Can Chinese companies fully catch up with U.S. frontier-model companies technologically?

🔹 Will the performance gap between open-weight and closed models disappear within the next few generations?

🔹 To what extent will independent evaluations confirm the superiority suggested by selected benchmarks from companies such as OpenAI and Anthropic?

🔹 Will Chinese companies continue their open-weight strategies even for their most capable models?

🔹 Can countries outside the U.S. and China re-enter the frontier-model race?

🚀 Is a new era beginning for open-weight LLMs?

It is still too early to say that open-weight models have completely overtaken closed systems.

But something else can now be said with much more confidence: open-weight models are no longer sitting at the edge of the frontier AI race. They are moving closer to its center.

The rapid appearance of Inkling, Kimi K3, and Qwen3.8 is one of the clearest signs of this shift.

If the performance gap between open and closed models continues to narrow, the consequences will extend far beyond benchmark tables. It will affect how AI infrastructure is built, which model providers companies depend on, how countries think about AI sovereignty, and even the future of today’s API-centered business models.

What do you think? When will open-weight models catch closed models in overall performance? Or will the most capable models always remain closed?


Murat Karakaya Academy

#MuratKarakayaAkademi #ArtificialIntelligence #LLM #OpenWeights #OpenSourceAI #KimiK3 #Qwen38 #Inkling #GPT56 #ClaudeFable5 #AgenticAI #GenerativeAI

```

ARTIFICIAL INTELLIGENCE • LLM • AGENTIC AI • 2027

ARTIFICIAL INTELLIGENCE • LLM • AGENTIC AI • 2027

Witnessing the AI Revolution: From Deep Learning to Long-Running AI Agents

Over the past 15 years, artificial intelligence has not merely produced better models. We have moved from systems that recognize images to models that converse with us, from conversational models to agents that use tools, and from there to AI systems capable of working on a problem continuously for hours or even days. So what will be the next major breakthrough in 2027?

A Question That Emerged from a Conversation

Today, while talking with a friend who is also deeply interested in artificial intelligence, our conversation naturally turned to the extraordinary pace of development in AI over the past few years.

After a while, we started recalling the technologies we had been discussing over the years: Deep Learning, RNNs, LSTMs, Transformers, large language models, multimodal models, reasoning models, RAG, function calling, agents, MCP, computer use, coding agents, and most recently, AI agents capable of working independently for extended periods of time.

When you put all of these developments side by side, the speed of change becomes much more visible.

In fact, we have not simply developed more capable AI models over the past few years. The way we use artificial intelligence itself has changed several times.

Artificial intelligence is evolving from a system that recognizes patterns into a system that talks to us, reasons about problems, uses tools, and increasingly performs complete tasks on its own.

🧠 1. Deep Learning: The Beginning of the Major Breakthrough

Artificial neural networks were certainly not new. However, for many years they could not demonstrate their full potential because of limitations in data availability, computing power, and training techniques.

In 2012, AlexNet’s success in the ImageNet competition demonstrated how powerful deep neural networks trained on GPUs could become.

The breakthrough that started in computer vision quickly spread to speech recognition, natural language processing, and many other machine learning problems.

From today’s perspective, this may seem almost ordinary, but one of the key foundations of the AI revolution we are experiencing today was established here:

Larger neural networks + more data + more computing power.

🔁 2. RNNs, LSTMs, and Sequence-to-Sequence Models

While CNN architectures achieved remarkable success in image processing, natural language presented a different challenge: language is sequential.

RNNs, and particularly LSTMs, made it possible to retain information from previous steps and process sentences, conversations, and time-series data more effectively.

Then came sequence-to-sequence, or seq2seq, architectures. They enabled a model to receive one sequence and generate another sequence as output. Machine translation became one of the most important applications of this idea.

Today, we consider it completely natural to give a model a piece of text and receive another piece of generated text in return. Yet encoder-decoder and seq2seq architectures were important milestones on the road toward modern generative AI.

👁️ 3. Attention: Teaching the Model Where to Look

There is an important chronological detail here: the attention mechanism did not begin with the Transformer.

Around 2014, particularly in neural machine translation research, models began to learn how to focus more heavily on relevant parts of the source sentence while generating each part of the output.

This idea was extremely important because a model no longer had to compress the entire input into a single fixed representation.

⚡ 4. 2017: Attention Is All You Need

In 2017, the publication of Attention Is All You Need became one of the most important turning points in the history of artificial intelligence.

The Transformer architecture greatly reduced the dependence on the sequential processing used by RNNs and placed self-attention at the center of the architecture.

With self-attention, relationships between different words in a text could be modeled directly, while computations could be parallelized far more efficiently.

Today, the architectural roots of major model families such as GPT, Claude, Gemini, Llama, Qwen, and Mistral can largely be traced back to this transformation.

📚 5. Pre-Training and the Foundation Model Paradigm

After the Transformer, another major paradigm gained momentum.

Instead of training a separate model from scratch for every task, researchers began training large general-purpose models on massive datasets and then adapting them to different downstream tasks.

Models such as BERT demonstrated how powerful this approach could be in natural language processing.

The discussion gradually moved beyond the question of “Which architecture performs better?” toward a broader question: “Can we train general-purpose models capable of performing many different tasks?”

📈 6. Scaling and GPT-3

Around 2020, the relationship between model size, training data, computational resources, and performance became far more visible.

Scaling law research suggested that larger models did not simply perform the same tasks slightly better. In some cases, new capabilities appeared as scale increased.

GPT-3 made another important phenomenon highly visible:

A model could perform many different tasks without being retrained for each one, simply by receiving instructions or a few examples in the prompt.

Zero-shot and few-shot prompting became increasingly important.

📄 7. RAG: Moving Beyond the Model’s Parameters

One of the major limitations of large language models was their inability to reliably access current, organization-specific, or previously unseen information.

Retrieval-Augmented Generation, or RAG, approached this problem from a different direction.

Instead of trying to encode every piece of information inside the model’s parameters, relevant information could be retrieved from an external knowledge source and added to the model’s context.

Over time, RAG evolved into more sophisticated architectures involving embeddings, vector databases, semantic search, hybrid search, reranking, query rewriting, and knowledge bases.

This approach remains extremely important today, particularly in enterprise AI applications.

🎨 8. Diffusion Models and Generative Visual AI

Meanwhile, the AI revolution was not limited to text.

With the rise of diffusion models, text-to-image generation rapidly became mainstream.

This was followed by major progress in image editing, style transfer, image-to-video generation, text-to-video generation, speech synthesis, and AI-generated music.

The concept of Generative AI was no longer limited to models that generated text.

👀 9. Multimodal Models

The next important step was bringing different types of data together within the same model.

The model was no longer limited to understanding the text we typed.

📷 It can understand photographs.
📊 It can interpret charts.
📑 It can read documents.
🎙️ It can understand audio.
🖥️ It can see a computer screen.
🎥 It can analyze video.

Because the real world is not composed only of text, multimodality became an essential step toward more general-purpose artificial intelligence.

💬 10. 2022: ChatGPT Brings AI to the Masses

The release of ChatGPT on November 30, 2022 was, in my view, not only a technological milestone but also a major social turning point.

For the first time, large language models entered the daily lives of ordinary users at massive scale.

Teachers, students, doctors, lawyers, engineers, managers, artists, software developers, and entrepreneurs all began interacting with the same technology.

Instruction tuning and alignment methods based on human feedback made interaction with these systems far more natural from the user’s perspective.

At the same time, a new term entered everyday vocabulary:

Prompt Engineering.

🔧 11. Function Calling and Tool Use

Large language models could generate highly impressive text, but their ability to interact with real-world systems was limited.

Function calling and tool use began to change that.

A model could now call an API, access a database, use a calculator, perform a web search, or invoke another software component.

Old approach: “Answer the question.”
New approach: “Use tools when necessary to solve the problem.”

This seemingly small change later became one of the foundations of agentic systems.

🧠 12. Reasoning Models

In 2024, the rise of reasoning models introduced another important shift.

The focus was no longer only on adding more parameters to the model. Models could also spend more computation on a difficult problem before producing an answer.

The concept of test-time compute became increasingly important.

Especially in mathematics, science, coding, and complex problem solving, the idea of the “larger model” was joined by another idea: the “model that spends more time reasoning about the problem.”

💻 13. From Chatbot to Coding Agent

In my view, one of the most important recent transformations has occurred in software development.

At first, we asked models to generate small pieces of code.

Then came copilot-style systems embedded directly inside IDEs.

Later, AI systems began reading entire repositories, creating files, running terminal commands, debugging errors, writing tests, and refactoring code.

💬 A system that talks about code
➡️ A system that suggests code
➡️ A system that writes code
➡️ A system that works across a repository
➡️ An agent that takes responsibility for a software development task

🔌 14. MCP: Standard Connections Between AI and the Outside World

As agent systems became more common, another problem became increasingly visible.

Every AI application had to be integrated separately with GitHub, file systems, databases, enterprise applications, and many other external services.

The Model Context Protocol, or MCP, became an important development for establishing more standardized connections between AI applications, tools, and data sources.

I do not think MCP should be viewed merely as a technical protocol. For an agent ecosystem to grow, models need standardized and manageable ways to connect to the outside world.

🧩 15. Agent Skills

The next step was to stop explaining the same procedures to an agent again and again, and instead define reusable skills that describe how particular tasks should be performed.

A skill can, for example, teach an agent:

📊 How to analyze an Excel file,

📄 How to prepare a professional report,

🧪 How to execute a testing procedure,

🏢 How to follow an organization-specific workflow,

💻 How to perform a particular software development task.

The agent can then use that skill whenever it becomes relevant.

This seems especially important in the transition from general-purpose AI toward specialized digital workers.

🖥️ 16. Computer Use and Browser Use

Agents then began doing more than simply calling APIs.

They started seeing screens, moving the mouse, clicking buttons, navigating websites, and interacting with the same graphical user interfaces that humans use.

I believe this is a critical development.

Most software in the world does not have an API specifically designed for AI. But almost every application has a user interface designed for humans.

With computer use, the user interface itself is effectively becoming another integration layer that AI systems can operate.

🌍 17. Open Models and Local AI

While all of these developments were taking place, another transformation was happening in parallel.

Open-weight models rapidly narrowed the gap with closed models across many practical use cases.

Model families such as Llama, Mistral, Qwen, and DeepSeek helped build a powerful open ecosystem.

Quantization, more efficient inference engines, and advances in hardware made it possible to run models locally that previously would have required large data centers.

Artificial intelligence no longer has to exist only as a cloud service accessed over the internet.

Local AI systems that keep organizational or personal data on-premise are becoming increasingly realistic alternatives.

The emergence of personal agent software that can run continuously on a local computer is another important part of this transformation.

🤖 18. 2026: The Year of AI Agents

For some time, many people had predicted that 2026 would become “the year of AI agents.”

I think that prediction has largely come true.

However, the meaning of the word “agent” has also changed.

Early agent architectures looked roughly like this:

LLM → call a tool → observe the result → call another tool if necessary → answer

Today, we are talking about far more comprehensive systems.

🧭 They can plan.

🧰 They can choose the right tool.

🗂️ They can manage files and context.

🧠 They can preserve state across a task.

🤝 They can delegate work to other agents.

💻 They can run code and terminal commands.

🌐 They can perform actions on the web.

🔍 They can evaluate their own outputs.

🛡️ They can operate within permission boundaries.

👤 They can request human approval when necessary.

🏗️ 19. The Next Step: Agent Harnesses

One concept I find particularly important in recent developments is the agent harness.

As we assign agents longer and more complicated tasks, the intelligence of the underlying model alone is no longer enough.

A model can perform extremely well during the first few steps and then drift in the wrong direction.

It can incorrectly assume that the task is complete.

It can forget previous decisions.

It can reach a conclusion without testing its work.

Its context can gradually degrade during long-running tasks.

This is why we increasingly need an execution environment around the model that manages and supervises its work.

📝 It preserves the plan.
🔄 It triggers replanning when necessary.
🧪 It tests intermediate and final results.
🔍 It checks for missing work.
📚 It manages context.
📈 It monitors progress.
✅ It can independently verify the output.
🛑 It helps determine when the agent should stop.

This moves us from chatbots that work for a few minutes toward AI systems that can work toward a goal for hours, or potentially much longer when supported by the right infrastructure.

I THINK THIS IS THE REAL BREAKTHROUGH
The competition is no longer simply about “who has the best LLM?”
The real competition is about how effectively a model can operate with the right data, the right tools, the right memory, the right skills, the right security mechanisms, and the right verification environment.

How Has the Short History of Modern AI Actually Changed?

🧠 2010s → The model recognizes patterns.
💬 2022 → The model talks to us.
👀 2023–2024 → The model sees and hears.
🧠 2024 → The model spends more time reasoning about the problem.
🔧 2023–2025 → The model uses tools.
🖥️ 2024–2025 → The model uses computers and browsers.
🤖 2025–2026 → The model becomes an agent.
🏗️ 2026 → The agent is placed inside an environment designed for long-running work.

🔮 So What Will Define AI in 2027?

I do not think the next major leap will necessarily be just another, slightly more powerful conversational model.

One of my strongest candidates is the persistent agent: a personal or enterprise AI system that can work continuously while preserving its state and context over time.

In today’s chatbot paradigm, we go to the AI.

We open a window, ask a question, receive an answer, and then continue working ourselves.

The next stage may be very different: AI could become a persistent part of our working environment.

🧠 It could recall our previous work when relevant.

📧 It could work with our email and messages.

📅 It could follow calendars and tasks.

📁 It could understand our files and projects.

🌐 It could conduct research when necessary.

💻 It could perform actions on our computers.

🧩 It could use the skills it needs.

🔌 It could connect to enterprise systems.

🤝 It could collaborate with specialized agents.

🛡️ It could understand its permission boundaries.

👤 It could request human approval for critical actions.

✅ It could verify its own work before reporting completion.

☁️ Cloud or Local?

I also do not think this future will be entirely cloud-based.

As personal and enterprise data becomes more important, hybrid architectures will probably become increasingly attractive.

Small and fast models may run continuously on local devices.

When more complex reasoning is required, the system could call more powerful cloud models.

Organization-specific knowledge could remain on-premise.

Highly sensitive tasks could be handled entirely by local models.

In the future, instead of asking “Which model are you using?” we may increasingly ask: “What kind of AI architecture is your system running on?”

🤖 Another Strong Candidate: Physical AI

Another major field to watch is artificial intelligence moving beyond the computer screen and into the physical world.

When vision, reasoning, planning, and agent technologies are combined with robotics, we may see major advances in embodied AI or physical AI.

An agent that learns how to click a button on a computer screen today and an AI system that controls a physical robot tomorrow may be more closely related problems than they first appear.

🚀 Perhaps the Most Important Part of the Revolution Is Just Beginning

When I look back at the developments of the past 15 years, the most interesting change for me is how our questions have changed.

During the Deep Learning era, we asked:

“Can the machine recognize this?”

During the LLM era, the question became:

“Can the machine understand and explain this?”

With reasoning models, we started asking:

“Can the machine solve this problem?”

In the age of agents, the question is very different:

“Can I give this entire task to AI from beginning to end?”

I do not think this is a small change.

For most of the history of computing, humans learned how to use computers.

Now computers are beginning to learn how to use the same tools that we use.

Perhaps the most important transformation of the next few years will happen precisely here.

🤔 What Is Your Prediction?

Which technology or approach do you think will become the defining AI trend of 2027?

🤖 Persistent personal AI agents?

🏢 Enterprise agent systems?

🏗️ Agent harnesses and long-running autonomous work?

🧠 More powerful reasoning models?

🖥️ Computer use and browser use?

🌍 Open and locally running models?

🤖 Physical AI and robotics?

Or will the next major breakthrough come from something that we are not yet talking about enough?

Selected Milestones

Krizhevsky, Sutskever & Hinton — ImageNet Classification with Deep Convolutional Neural Networks, 2012

Sutskever, Vinyals & Le — Sequence to Sequence Learning with Neural Networks, 2014

Bahdanau, Cho & Bengio — Neural Machine Translation by Jointly Learning to Align and Translate, 2014

Vaswani et al. — Attention Is All You Need, 2017

Devlin et al. — BERT, 2018

Kaplan et al. — Scaling Laws for Neural Language Models, 2020

Brown et al. — Language Models are Few-Shot Learners, 2020

Lewis et al. — Retrieval-Augmented Generation, 2020

Ho, Jain & Abbeel — Denoising Diffusion Probabilistic Models, 2020

#MuratKarakayaAkademi #ArtificialIntelligence   #GenerativeAI   #LLM   #AIAgents   #AgenticAI   #DeepLearning   #Transformer   #ReasoningModels   #RAG   #MCP   #AgentSkills   #ComputerUse   #OpenSourceAI   #LocalAI   #AIEngineering   #FutureOfAI

Artificial Intelligence • R&D • Project Management

Artificial Intelligence • R&D • Project Management
```

🚨 Why Do AI Projects Fail?

Project proposals are becoming increasingly impressive. But can we say the same about the engineering capacity of the teams expected to turn those proposals into working systems?

```

Recently, I have been noticing the same problem more and more often during project evaluations.

When you start reading an AI-related project proposal, you often see almost every recent method, technique, and technology somehow included in the project:

Large language models, RAG, AI agents, vector databases, personalization, fine-tuning, multi-agent architectures, real-time analytics, prediction systems, decision support, explainable AI...

On paper, the project looks highly impressive.

However, once you meet the team face to face and begin asking more detailed questions, a very different picture often emerges.

👥 The technology is in the proposal — but is the experience in the team?

Sometimes, the engineers on the project team have never developed even a small-scale application using the technologies described in the proposal.

The company itself may also have no prior experience building, deploying, operating, or maintaining products based on those methods.

After a few questions, the following issues begin to surface:

  • Why was this technology selected?
  • What real problem does it solve in the project?
  • Why is it being used together with the other components?
  • Were alternative methods evaluated?
  • How will success be measured?
  • If the system fails, how will the team determine which component failed and why?

Unfortunately, in some projects, there are no clear answers.

The technology choices appear to be less the result of an engineering analysis and more the result of connecting fashionable concepts to one another.

💡 The problem is not that too many technologies are being used. The problem is that the team cannot explain, from an engineering perspective, why each technology is actually necessary.

🧩 A list of technologies is not a system design

Today, it is relatively easy to produce impressive project documents with generative AI.

You can list the advantages of every technology. You can summarize their benefits from the literature. You can draw an architecture with a few boxes and arrows.

But none of this, by itself, constitutes systems engineering.

In a real system design, every component has a reason to exist.

A component should not be added because it is fashionable. It should be added because it satisfies a specific requirement.

For every selected method, the team should be able to explain:

  • which requirement it addresses,
  • which alternatives it was compared against,
  • which metric it is expected to improve,
  • which costs it introduces,
  • and which new risks it adds to the system.

Otherwise, what we have is not a system architecture. It is merely a list of current technologies.

🔍 Saying “we will use RAG” is easy. Building a working RAG system is not.

For example, the following sentence has become very common in AI project proposals:

“A RAG-based artificial intelligence system will be developed to generate reliable answers from corporate documents.”

Fine.

But the real engineering problem begins exactly where that sentence ends.

📌 How will system performance be measured?

📌 How will the benchmark dataset be created?

📌 How will real user questions be represented in the evaluation set?

📌 How will retrieval quality be separated from final answer quality?

📌 How will chunk size and overlap be selected?

📌 Which experiment will determine the embedding model?

📌 Which metadata fields will distinguish text chunks from tens of thousands of documents?

📌 If multiple versions of the same document exist, which one will be used?

📌 How will obsolete or withdrawn information be removed from the system?

📌 How will unauthorized users be prevented from retrieving restricted content?

📌 How will hallucination be measured?

📌 How should the system behave when a reliable answer cannot be generated?

📌 How will latency be controlled under thousands of concurrent users?

📌 How will the trade-off between cost, accuracy, and response time be managed?

Every one of these is a real-world engineering problem.

Yet in many project documents, instead of answering these questions, the proposal simply explains the generic benefits of RAG.

That tells us that the project may be conceptually defined, but its operational reality has not been sufficiently considered.

⚠️ Using AI to prepare a project is not the same as developing an AI project

I believe a new risk has emerged in recent years.

Generative AI has made project preparation dramatically easier.

Literature reviews, technology descriptions, methodology sections, and technical documentation that once took weeks can now be produced much faster.

This is not inherently a bad development.

On the contrary, when used properly, it can create enormous productivity gains.

The problem begins when the process turns into something like this:

🤖 Ask AI to write the project proposal.

💰 Apply for funding.

🧑‍💻 If the funding is approved, ask AI to write the code too.

This mindset can easily turn into a dangerous form of shortcut-driven project development.

The engineering gap between a project proposal and a real system cannot automatically be closed simply by using more AI.

🧠 AI does not automatically create experience that the team does not have

Today’s generative AI systems can write code, suggest architectures, create APIs, generate test scenarios, and produce documentation.

But there is still a critical distinction:

AI can generate a solution for you. But determining whether that solution is actually the right solution for your problem still requires knowledge and experience.

Because real engineering is not simply about generating code.

Real engineering means:

  • defining the right problem,
  • identifying the requirements,
  • comparing alternative approaches,
  • defining measurable success criteria,
  • anticipating bottlenecks,
  • handling unexpected cases,
  • testing the system under real user behavior,
  • and considering scalability and sustainability.

📉 So what happens in the end?

A system that looked extremely strong in the project proposal fails to deliver the expected performance during implementation.

At first, a few demonstrations are prepared.

The system performs well on carefully selected examples.

But once real data, real users, and real operational conditions enter the picture, problems begin to appear.

Costs increase.

Latency rises.

Data quality problems emerge.

Authorization requirements become important.

The model begins producing unexpected answers.

Real user questions do not resemble the clean examples used in demonstrations.

And because there is no properly designed benchmark, the team may not even know which change actually improves the system.

As time passes, the project schedule slips.

The team changes.

Employees leave.

New team members do not understand why previous architectural decisions were made.

Eventually, the project is either scaled down, diverted from its original objective, or left behind as an unused prototype.

🎯 Perhaps we are asking the wrong question

When evaluating AI projects, we often focus on this question:

“Which technologies will be used in the project?”

I believe there is a more important question:

“Does this team actually have the engineering capacity to build the system described in the proposal?”

This should be followed by several additional questions:

  • Has the team solved a similar problem before?
  • Has it at least built a small-scale prototype?
  • Does the team understand the technical risks?
  • Does it know which assumptions may turn out to be wrong?
  • Can it design experiments to measure success?
  • Does it know how to validate the system under real user load?

🏗️ How should a good AI project begin?

In my view, the process should begin with the problem, not with technology names.

Wrong starting point A better starting point
“Let’s use RAG.” “Let’s understand which information users cannot access and why.”
“Let’s add agents.” “Let’s identify which tasks genuinely require multi-step reasoning and decision-making.”
“Let’s personalize everything with an LLM.” “Let’s measure which user information actually improves response quality.”
“Let’s use the most powerful model.” “Let’s experimentally determine the right trade-off between quality, cost, and latency.”

🚀 A strong project proposal does not necessarily mean a strong project

Today, it is possible to prepare an excellent project proposal with the help of AI.

In fact, compared with just a few years ago, proposals can now be more professional, more technical, and far more persuasive.

But the quality of the proposal and the feasibility of the project are not the same thing.

AI can accelerate project preparation.

It can accelerate software development.

It can simplify documentation.

It can generate test scenarios.

It can suggest architectural alternatives.

But it does not magically create engineering experience, domain knowledge, systems thinking, or product development culture that the team does not already possess.

🎯 One of the biggest risks in AI projects may no longer be lack of access to technology. It may be the opposite: because access has become so easy, teams may begin highly ambitious projects without sufficient preparation or experience.

At the end of the day, we still need people who understand the problem, ask the right questions, design meaningful measurements, anticipate failure points, and architect the system from end to end.

💬 Have you also encountered AI projects that look extremely strong in project documentation but fail to translate into real-world systems?

#MuratKarakayaAkademi #ArtificialIntelligence #AI #GenerativeAI #LLM #RAG #AgenticAI #Engineering #SystemsEngineering #ProjectManagement #ResearchAndDevelopment #ProductDevelopment #AIProjects #Innovation #Technology