Blog/News·September 21, 2026

From AI transcriptions to AI doing work in meetings

What changes when the AI in the room sees what you see, remembers across every meeting, and does the work that follows.

Today: how every AI work tool treats its agent

You hire the smartest person on the planet. Near-photographic memory, endless patience, works around the clock.

Then you dial them into a few of your meetings over a crackling phone line. No screen, so they can't see the deck, the dashboard, or the document everyone else is staring at. You forbid them to remember anything beyond the words spoken in that one call. Half the time they're guessing who's even talking.

And then the whole industry wonders why the work doesn't get done properly.

Tomorrow: the agent is in the room

A real participant in work and meetings. Present where the work happens. Remembering across every meeting it has access to. Aware of what the rest of the organization is doing. With a clear understanding of roles and decision-making authority. Able to do the work discussed in the meeting, like any other colleague.

It sees what we see. It knows what we know. It does what we do.

Abstract

Meetings don't fail to produce work because the AI in them isn't smart enough. They fail because it's starved of context.

Context is more than a transcript. It's what was meant and why, who decided, what it changes, and what it now permits. It's the screen everyone was looking at, the meeting last week that settled the question, and the standup two teams over where somebody is about to do your work again. A transcript carries none of it.

Give an agent that context and it can do the work that follows the meeting: file the issue, draft the contract, update the CRM, open the pull request, before anyone leaves the call. Research on multi-agent systems says the same thing from the other side. Nearly four in five agent failures are specification and coordination problems that a stronger model leaves untouched. Agents don't need a bigger brain. They need an organization: a record of what was decided, a role that says what they may do, and someone with the authority to say no.

Three things let a team trust the result. Every output traces back to a human decision. Nothing enters the shared record without a human. And the system is judged on what got done, not on how much it produced.

The record, shared across every team and every agent, is a company hive mind: one shared context your people and your agents work from. It was too expensive to build until we brought the cost of looking at a whole meeting down 770 times. The hive mind remembers. It doesn't decide. What stays with people is judgement, and the meetings where judgement happens.

1. The gap between talking about work and doing it

Every team has a decision it has made more than once. A discussion that keeps resurfacing. An initiative suggested multiple times. A group of people think hard about something, reach a conclusion, and then the conclusion fails to influence the future the way it should. Work, and meetings especially, is a series of exploratory deep dives that arrive somewhere and typically leave nothing behind. So we repeat the same discussions, and complain that these meetings stop us from getting actual work done.

Wasted and duplicated work is bad within a team. It's even worse across them. Discussions are siloed, so we remain oblivious to all the other teams thinking hard about the same problems.

Atlassian's State of Teams 2026 puts a number on the wreckage. Seventy-five percent of employees observe work being duplicated. Eighty-seven percent say they lack the time or capacity to coordinate. The resulting breakdown costs the Fortune 500 an estimated $161 billion a year.

The breakdown is organizational. And currently AI amplifies it. More discussions, in more places, with even less coordination.

Three statistics from Atlassian's State of Teams 2026: 75% of employees watch the same work get done twice, 87% say they have no time or capacity left to coordinate, and $161B in annual cost across the Fortune 500 alone.

The obvious response is to write more down. Which does less than you'd hope. Software has already run the experiment. Every pull request in a codebase is already there. The entire story of why the code looks the way it does, every argument, every rejected alternative, is sitting in the history. Yet onboarding a new engineer is still brutally hard. The information exists and it's still not usable, because reading a thousand decisions in sequence isn't the same as understanding what the team was actually trying to do.

Having the history and being able to enter the conversation are different problems. Most tools solved the first one and called it done.

AI note-takers were the industry's answer to coordination. And they solved a real pain: the meeting gets written down without anyone taking minutes. It's an improvement, but it's not enough. Most meetings are still their own island, with no memory of the one before and no idea a related one is happening two teams over. Note-takers are a small example of a larger pattern. AI got pointed at the part of the problem that was easy to automate, and the part that actually costs money went untouched. MIT Media Lab's NANDA initiative found that 95% of enterprise AI pilots produce no measurable financial return, and named the cause organizational rather than technological.

95% of enterprise AI pilots produce no measurable financial return, shown as a grid of twenty circles with nineteen filled.

The transcript gets filed. Yet the next meeting starts from scratch.

Why meetings? And why now?

There's a reason we care about meetings specifically, and it's not (only) because meetings are annoying.

As AI takes over execution, the work left for humans is deciding, aligning, and communicating. That work happens in meetings, between people and increasingly between people and AI. Which makes the meeting the most valuable surface in the company.

As execution moves to the machines, we believe meetings are where the next decade of work lives.

Three companies own that surface today. Google, Microsoft and Zoom. Look at what they're building and how carefully they keep everyone else out of their platforms. The reason isn't video quality. It's context. If context is everything, then meetings are the largest context source any organization has, and whoever owns the meeting owns the context, aka the basis for all downstream work.

Whoever connects the meetings has something bigger: a company hive mind, one shared context that people and agents work from.

And what do the three of them give you from that surface? A transcript (if you're lucky).

2. A transcript is not context

A transcript records what was said. Context is what was meant and why, who decided, what it changes, and what it now permits.

A text-only record loses key information. The first loss is vision. An agent that can only read is working with one sense, and not the strongest one. If it sees the developer console and the error sitting in it, it's far more likely to understand the fix than if it hears someone try to describe the bug out loud. Think about what an agent needs the way you would think about what a person needs: to make a good decision about a situation, you generally have to look at it.

Vision is where large parts of technical meetings happen: the dashboard nobody reads out loud, the failing test, the design that takes 15 minutes to describe poorly and one glance to understand. Transcript-only tools have none of it. What they preserve is the commentary on the thing, with the thing removed.

Figure 2.1
Concentric rings: this meeting at the centre, then the meeting series, then related meetings elsewhere, then the organization. A transcript reaches only the innermost circle.
Source: Tana

The second loss is structure. A transcript is linear and decisions are not. Reading one back gives you the order things were said in, which is close to useless, because the important sentence is rarely the last one and the person who settled the question is rarely the person who talked most.

What an agent needs is the decision trace: who decided, when, on what basis, and what that decision now permits or forbids. Every decision carries a constraint on everything downstream of it, and an agent blind to that constraint will cheerfully propose the thing you ruled out in March. Or implement the opposing idea your summer intern pitched during the all-hands.

The meeting in front of you is only the innermost ring. Beyond it sits the last meeting in the series, the next one, and what that next one will have to deal with. Beyond that sit the meetings you were never in, happening elsewhere this week, bearing directly on what you're deciding.

An agent without vision works with a fraction of the relevant context. One that sees the room does better. One that also knows the series, the related meetings, and what the rest of the company is doing can do the work: draft the ticket, open the pull request, write the spec, before anyone leaves the call. It knows what was ruled out in March, so the room doesn't decide it twice. It sees the other team's standup, so two teams don't build the same thing.

Say the meeting is a product status review. A transcript-based tool can, at best, remind you what was said in the last product status review. An agent with organizational context can reach into the product groups underneath, roll their current state up into the meeting, and put it in front of you while the meeting is still running. That material was never in a transcript.

The everyday version happens each morning. You say what you're planning to work on. Another team has already done it, or is doing it right now, and said so in their own standup. If standups are visible across the company, the agent catches the duplication in the moment, while it still costs nothing to fix. This is the company hive mind at work.

The right mental model for what we're trying to build isn't a note-taker or recorder. It's an extremely talented chief of staff: someone who went to every meeting, remembers it all, and can surface the thing from three weeks ago in another department that changes what you're about to decide.

A transcript also can't mark which parts of a conversation were meant to survive the week and which were thinking out loud. Google Wave (launched May 2010) had great promise and exactly this problem: everything was preserved with equal weight, which is the same as preserving nothing. Without a record of why, every decision gets defended again the next time it comes up.

Nobody built this before because the curation job was too much work. Wave could turn a chat into an artifact, but someone had to play secretary and pull things out by hand, so nobody did. The job always existed and was always unpaid. An agent will do it, at volume, without getting bored.

A transcript tells you what was said. An agent that knows what it meant, who decided it, and what the rest of the company is doing can act on it, accurately, before anyone leaves the call.

3. Agents don't need a bigger brain. They need an organization.

Most agents that try to act still fail. This is what it takes to build one you'd trust with the work.

When an agent fails, everyone's instinct is to reach for a better model. That's usually the wrong lever. The failures are organizational, and you can't fix an organizational problem with a smarter model. What the agent is missing is what a new hire is missing on day one: a record of what's already been decided, a role that says what it's allowed to do, and someone with the authority to tell it no.

The record, the decisions with their reasoning attached, is covered. This chapter gives the agent a role and someone who can say no.

Researchers at UC Berkeley annotated more than two hundred traces of multi-agent systems failing, across seven different frameworks, and sorted every failure into fourteen modes. Roughly 42% were specification problems. Another 37% were agents talking past each other. The remaining 21% were verification failures, and those are a design problem too: the checks ran, they just only looked at the surface. Their own conclusion is worth quoting: "improvements in the base model capabilities will be insufficient," because "good MAS [multi-agent system] design requires organizational understanding."

The individual failure modes are telling. The three most common were step repetition, at 17.1%, agents doing work that had already been done. Reasoning-action mismatch, at 14.0%, agents saying one thing and doing another. And failing to ask for clarification, at 11.7%, agents guessing at what was wanted rather than checking. Not one of those is a knowledge problem. They're failures of process: no record of what had been done, no check that action matched intent, no way to raise a hand. Every one of them would be a management problem if a person were doing it.

Figure 3.1
Where multi-agent systems fail: 41.8% specification issues, 36.9% inter-agent misalignment, 21.3% task verification, with the five most common individual failure modes below.

Here's what that looks like up close. An agent was asked to build a locking system. The design it wrote was sound: a table of who holds which lock, and subscribers that update the table whenever a lock changes hands. Every piece the spec asked for was there and correctly named. But the implementation didn't do what the design said. Instead of updating the table when something changed, it deleted the whole table every four seconds and rebuilt it from scratch.

The tests passed. At test scale, brute force gives the same answers as the real design, so nothing looked wrong. Under real load it would have fallen over, and nobody would have known why, because the documentation described an incremental system that didn't exist.

This is the reasoning-action mismatch from the Berkeley study, inside a single agent. The model wasn't confused. It wrote a correct design and then cut a corner nobody was checking. A smarter model doesn't fix that. A role does: someone whose job is to open the implementation and hold it against the description, and who can block it when they don't match.

That was one decision, caught because someone happened to look. An agent working at real volume makes a thousand like it a day, and nobody looks at any of them.

The fix is a structure that makes a thousand small decisions reviewable, which is a problem human institutions solved a long time ago and gave an unfashionable name.

Bureaucracy. It's worth using the word deliberately rather than reaching for a softer one, because the softer ones have all been used up by software that didn't mean them.

Look at what a bureaucracy is if you strip the connotation. Forms are the data model: they define what has to be true before a case can move. Role definitions are the instructions: who handles what, and what they're permitted to decide. Escalation paths are the exception handling. The whole apparatus exists so that a similar case gets processed in a similar way by different people on different days, which is a hard problem, and it was solved well enough to run Rome's imperial civil service for centuries, compartmentalize the Manhattan Project, and put a checklist in the cockpit of every aircraft you've ever flown on.

A machine made of people.

Its bad reputation is earned. It was designed for low bandwidth and high latency, for a world where a form took three days to reach the next desk, and once you optimize for zero variance you also optimize for zero adaptation. We don't wish to keep that part.

Bureaucracy got a bad name because it ran on paper. The parts that made it work were roles, records, and the right to say no.

With agents, it looks like this. Team lead agents hold the policies. Specialist agents go deeper on one vertical than their lead does. Judge agents rule on whether work meets the standard. The rule for when to split one agent into two is simple: as long as there's no conflict of interest, give one agent more skills. The moment you need something to say no, separate it out. And separate it out with teeth, because this is where most review layers fail.

A reviewer that can only comment is decoration. A reviewer that can block is a control.

Figure 3.2
Agent topology in three layers: a team lead agent holding the policies, three specialist agents below it, and a judge agent that can block work from shipping.
Source: Tana, after Olav Sindre Kriken

Look at what that structure is made of. A record of what was decided. Roles that say who may do what. Something with the authority to block. None of it comes from the model vendor. All of it comes from your own organization, and most of it you already produce every week, in the meetings where you decide things, if anything is keeping it.

The frontier labs are building the one true brain. Everyone else is building the society that manages the smaller brains the labs give them access to. Both are serious engineering. Only one of them is available to you.

You're not going to out-model the labs. You can out-organize everyone who's trying to.

4. What doing work in meetings actually means

Give the agent a record, a role, and someone who can say no, and it can do the work discussed in meetings.

"Doing work in meetings" isn't a phrase about engineering teams. A lawyer meets a client and owes them a revised agreement. A salesperson runs a call and owes a follow-up, a contract with the new terms in it, and a CRM record that matches what was said. HR runs a calibration meeting and owes every manager a written rationale by Friday. A manager runs a one-on-one and owes somebody a decision they promised to make. Every one of those meetings ends with work that a human then goes away and does. This is about to change.

Once the room has decided, everything that follows from the decision is the agent's job.

Not the judgement. Not the decision. The part afterwards where a person retypes into four systems what everyone in the room already knew twenty minutes ago.

Figure 4.1
Four disciplines and what each owes after a meeting: sales, legal, product and engineering, and people and management.
Source: Tana

Then it keeps going, which is the part most tools never reach. Take customer onboarding. The customer struggles with something during the call. That gets filed immediately as a product improvement rather than as a line in a summary nobody reads. Agents pick it up, dig into what the underlying problem was, and propose fixes. On a good day they open a pull request, so the same issue can be looked at from the product side and the engineering side at once. Same chain when the customer finds a bug.

The full distance: something said out loud in a meeting, and a reviewed change sitting in the repository, with the line connecting them intact.

Four bugs filed in Linear before the standup ended beats any productivity claim you could write.

Figure 4.2
Three screens from one standup: the issue pinned in the meeting side panel, the filed issue record, and the merged pull request.
Source: Tana. Three screens out of one 40-minute standup, August 2026.

5. Three things let teams trust work agents produce

We believe the following three things are required for teams to trust work done by agents at scale.

1. Every output traces back to a human decision

A plausible output and a grounded one look identical until somebody checks. That's the problem. A team can run for months on outputs nobody verified, and the cost shows up only when a decision gets made on something the model inferred rather than something a person said.

So every consequential output carries its provenance: the point in the meeting where the claim was made or the decision taken, and who made it. This is the decision trace put to work: who decided, when, on what basis. The record exists so that anything built on it can point back to it.

The provenance must answer one question: did a person decide this, or did the model? That's what a reviewer with the authority to block needs. A judge can't block what it can't trace.

An output nobody can trace to a decision is an assertion wearing a citation.

Figure 5.1
An issue filed from a design review. The record names the page that was on screen and the meeting it was discussed in, links the related issue, and attaches the annotated screenshot that prompted it.
Source: Tana

2. Nothing enters the shared record without a human

The agent proposes a change. A person reviews it and decides whether to merge it. The agent can't update the company source of truth without human approval.

This protects the first point. Outputs are trustworthy because they trace back to human decisions in the record. If an agent can write to that record unsupervised, an output can trace back to a decision the model invented, and the trace proves nothing. Keeping a person at the merge keeps the record human, so the trail behind every output ends at a person.

This is the reviewer with the authority to block, turned into a product rule. For the shared record, the judge is a person in a position to catch mistakes before AI amplifies them. Remove that person and you remove the only one who could be blamed. Nobody is accountable for a record they never approved, so the slop becomes the model's fault, which means it becomes nobody's.

3. Judge the system on what got done, not on how much it produced

Output volume is the easiest thing to count and the least useful. Optimize for it and the agent produces more of what nobody asked for.

The useful measures come from the record. Did committed actions get done? Did decisions stick, or get remade? To answer these you need a record that holds the commitment and the decision. That's what the record is for, and the first two points are what make it trustworthy: outputs trace back to it, and people control what goes in. The record acts as the source of truth and the scoreboard.

This has an uncomfortable second-order effect. Today, status travels upward through people. A team lead writes the update, a director compresses five updates into one slide, and each hand-off is a chance to decide what counts as on track and which miss goes unmentioned. The report is the author's version of the work, and writing a good version is a skill companies have historically rewarded.

A record of what happened removes the author. If the top-level view assembles from commitments made in meetings and whether they were kept, no summary slide is needed. This is the company hive mind seen from above. A missed commitment shows up on its own, before the person who missed it can decide how to frame it. People will resist. Especially those good at their jobs, who have been rewarded for managing the story.

But the shift is worth it. A company that reads its state from the record sees where work is stuck while it's cheap to fix. A company that reads it from prepared slides finds out one reporting cycle later and pays more.

Everything in this paper so far is design around the model, not changing the model itself. Two studies say that's where the return comes from.

MIT looked at why 95% of enterprise AI pilots returned nothing. The ones that did return shared three properties: a narrow scope, memory that carried between sessions, and a feedback loop that let the system correct what it got wrong. Scope is the role. Memory is the record. The feedback loop is the reviewer who can block. Model choice wasn't the deciding factor.

A study from Berkeley is even sharper. Holding the model constant and changing only the design of the system around it raised task success by up to 15.6%. No new capability, no larger model, no more expensive inference. The same models, organized better.

Figure 5.2
Five fields a decision needs to be accountable: source, reasoning, decision, owner and outcome, each with what breaks without it and how to check your own system has it.
Source: Tana

Trust in agent work comes from a human-controlled record that every output points to and every outcome is measured against.

6. Why everyone stopped at the transcript

Earlier we said nobody did the curation because it was unpaid work. That was true until agents could do it. Then it was true for a different reason: doing it properly cost more than anyone would pay.

Everything in this paper sounds like it should cost a fortune. An agent that sees a whole meeting, remembers across a series, and reaches into the rest of the organization is doing orders of magnitude more work than one that reads a transcript. The market stopped because it's hard to do while keeping it economically feasible, and it stays hard. Not hard as in one clever idea nobody had. Hard as in years of engineering effort, applied continuously, against a problem that gets more expensive every time the product gets better. Most teams looked at that and made the sensible call: capture the audio, transcribe it, ship.

We did the foolish thing. Not because it was easy, because we thought it would be easy. It wasn't. But sometimes foolishness pays off. On the same meeting, processed end to end, our cost came down by a factor of 770. Not 770%. 770 times.

770 times less, for the same meeting processed end to end. A bar running the full width of the figure is what one meeting used to cost; what it costs now is a two-pixel sliver at its left edge, drawn to the same scale.

That number is why we can afford to look at everything while a competitor has to choose which context to ignore. A hive mind that skims isn't one.

This affects you even if you never see a bill. Cost discipline isn't housekeeping. It decides what the product is allowed to do for you.

A vendor that hasn't solved cost efficiency will quietly cap how much of your meeting it pays attention to. It will look less often. It will handle the first hour properly and skim the rest. And it will call these limits focus, or noise reduction.

Technical block

Treat your cost data as a research corpus.

Point a coding agent at your own usage data and have it produce the breakdown. Run a second job over that breakdown and your codebase, asking it to find every optimization available and rank them by return, effort and risk. Implement, measure, repeat. Run it as a continuous loop and it keeps compounding.

Figure 6.2
A five-step loop: cost data, breakdown, ranked fixes by return effort and risk, implement, measure, then back to the start.
Source: Tana

Which brings us back to the transcript. Considering the full context stopped being a budget question and became an engineering one. That's what let us move past the transcript and build the record.

7. What stays human

AI can now do the work that follows the meeting. So what remains for us?

Your job becomes replacing yourself. You work at the automation frontier: the line between what the system can handle on its own and what still needs you. You train and oversee it until it works, then move to the next thing that doesn't. The record makes the frontier visible. It shows which of your commitments the system now closes without you, and which still come back to your desk.

This might sound like a threat, but it's what every senior person already does. The work you were doing three years ago isn't the work you're doing now, because you automated or delegated it. What changes is the speed.

Figure 7.1
The automation frontier as a line across the figure. Above it, your job: deciding what the standard is, owning what enters the record, ethical and value alignment, shared risk and the meetings about why. Below it, what AI can do: scheduling the follow-up meeting and catching duplicate work sit nearest the line, with filing the issue, opening the pull request, updating the CRM and drafting the contract further below.
Source: Tana

Never open a chat window to get an AI to do something. Open it to teach it how to do something.

The difference compounds. The first gets you an answer and leaves you exactly where you started, ready to ask again next week. The second gets you a capability: a decision in the record, with its reasoning attached, that the system handles alone next time. And the time after that. Most people are still using the most capable tool they have ever had as a faster way to produce one-off outputs.

Stay in ask-mode and the ratio between your time and the machine's will eat your day. You spend twenty minutes briefing, it works for thirty, and you're back. And again. And again. At some point you're an air traffic controller with no stretch of time long enough to think.

If you're only unblocking the machine, you've been demoted, not promoted.

The judge agents exist to absorb exactly this volume. Checking a thousand small decisions is their job. Deciding what the standard is, and owning what enters the record, is yours. If you find yourself checking every action, you've taken back a job the structure was supposed to take from you.

A system that's always available, always agreeable and never tired is comfortable in a way people are not. Agents will disagree with you about whether the work meets the standard. Only people will disagree with you about whether the standard is right. As we delegate easily verifiable tasks, the harder decisions and judgement calls remain. This is why we expect meetings to remain human work. The hive mind remembers. It doesn't decide.

What's left needs bodies present. Ethical and value alignment. Shared risk. Looking at each other and saying we're in this together. Debriefing the failures the record has already made impossible to hide.

The meeting is the most valuable surface in the company. Not because there will be more of them. There will be fewer, and the ones that survive are about why rather than what.

8. Start with one meeting

Every team has a decision it has made more than once, and a meeting that ends with everyone going back to their desks to do the work. Both have the same cause. The meeting left nothing behind that a person or an agent could act on.

That's what changes, and it changes one meeting at a time.

Start where the deciding happens. Pick one recurring meeting that already owes somebody a deliverable. The customer interview that becomes a research summary. The design review that becomes tickets. The sales call that becomes a contract and a CRM record. The sprint planning that becomes a backlog. The payoff is largest there and the risk is smallest, because the work is going to be done either way. You're only changing when, and by whom.

On the first run, the deliverable gets done before the call ends. The ticket is filed with the screen that prompted it. The contract carries the terms the room agreed. That's the visible part, and it's the part everyone expected.

After a few runs, something else has happened. The decisions from those meetings are in the record, with the reasoning attached and the person who made them. The agent stops asking what you meant, because it can look. The question you settled in the second meeting doesn't come back in the fifth. You have stopped briefing it and started teaching it, without doing anything except letting the record fill.

When a second team does the same, the record crosses the team boundary. Their standup is visible to your agent and yours to theirs. That's the company hive mind: one shared context your people and your agents work from.

The work two teams were about to do twice gets caught while it still costs nothing to stop.

Three things to know going in. Pick a meeting where something actually gets decided; a meeting that was never going to produce anything will teach the record nothing. Let the agent see the screen, not just hear the room; most of what matters in a working meeting is never said out loud. And give it more than one run. The record is what compounds, and it's empty on day one.

What you're building is delegation. To agents, which is the part everyone talks about. And to the people around you, who can act on a decision they weren't in the room for because the reasoning arrived with it. You stop briefing. You stop supervising line by line. What you get back is the stretch of time the good ideas needed.

Today, the transcript gets filed and the next meeting starts from scratch. Run one meeting in Tana and the next one starts where this one left off. Start with one meeting.

References

E1. Atlassian, State of Teams 2026. atlassian.com/blog/state-of-teams-2026

E2. Challapally, Pease, Raskar and Chari, The GenAI Divide: State of AI in Business 2025, MIT Media Lab Project NANDA, July 2025. Cited as reported; no stable first-party URL was available at the time of writing.

E3. Cemri et al., Why Do Multi-Agent LLM Systems Fail?, UC Berkeley, NeurIPS 2025. arXiv:2503.13657 and the MAST dataset

From AI transcriptions to AI doing work in meetings - Tana