A bet on a smaller LLM model that actually knows Polish law
A bet on a smaller LLM model that actually knows Polish law
A bet on a smaller LLM model that actually knows Polish law
Timeline
Sep 2025 – Sep 2026
Timeline
Sep 2025 – Sep 2026
Stage
Startup
Stage
Startup
Industry
Legal Ai
Industry
Legal Ai

My impact
Introduced a reactivation flow ("remind me about Gaius") alongside subscription-cancellation, used by 1 in 4 clients
Converting a multitude of unintuitive features into a cohesive application
Identified friction points in the core user journey
Increased the first-message rate by +31.6% absolute improvement by new signups in 4 months
My impact
Introduced a reactivation flow ("remind me about Gaius") alongside subscription-cancellation, used by 1 in 4 clients
Converting a multitude of unintuitive features into a cohesive application
Identified friction points in the core user journey
Increased the first-message rate by +31.6% absolute improvement by new signups in 4 months
context
At first, I thought AI legal agent would compete with other legal companies. Oh, how wrong I was, as it soon turned out that every lawyer was using GPT, Gemini or Perplexity.
So I asked myself:
what could we do better than big tech?
context
At first, I thought AI legal agent would compete with other legal companies. Oh, how wrong I was, as it soon turned out that every lawyer was using GPT, Gemini or Perplexity.
So I asked myself:
what could we do better than big tech?


goal
In Sep 2025, Gaius had a team of five and a major problem with usability.
App’s language was very technical
The visual hierarchy was poor
and we didn’t know exactly what kind of person we serve
but
the value was working 🎉
50 law firms were paying to find the right supreme court case reference, a source based on their line of case law or similar facts. And that's something I would call a great beginning.
After couple of weeks of talking with clients (not only about sales) I knew exactly with what expectations they came:
Quicker legal research without hallucination
Upload documents with sensitive data without worry
“proper” Ai answers
Wait, what 'proper Ai' answer even means?
Of course it depends and testing those assumptions in ongoing process.
goal
In Sep 2025, Gaius had a team of five and a major problem with usability.
App’s language was very technical
The visual hierarchy was poor
and we didn’t know exactly what kind of person we serve
but
the value was working 🎉
50 law firms were paying to find the right supreme court case reference, a source based on their line of case law or similar facts. And that's something I would call a great beginning.
After couple of weeks of talking with clients (not only about sales) I knew exactly with what expectations they came:
Quicker legal research without hallucination
Upload documents with sensitive data without worry
“proper” Ai answers
Wait, what 'proper Ai' answer even means?
Of course it depends and testing those assumptions in ongoing process.


strategy
Ran small-scale tests instead of shipping full features
In some cases, I run quick experiments before investing valuable development time. One example was displaying Lex token usage for each feature. The cost varied depending on the number of documents analysed and the type of task performed. More importantly, we couldn't accurately estimate token usage until the task had been completed.
The team started building a feature that would allow users to set a maximum token limit, but I convinced them to valide the need first.
User interviews and a smoke test, which displayed token usage after every response alongside a "We're working on it" message, revealed that users had little interest in monitoring token consumption. They wanted to know when they run out.
That allowed me to cut 20% of backlog saving our precious time.
strategy
Ran small-scale tests instead of shipping full features
In some cases, I run quick experiments before investing valuable development time. One example was displaying Lex token usage for each feature. The cost varied depending on the number of documents analysed and the type of task performed. More importantly, we couldn't accurately estimate token usage until the task had been completed.
The team started building a feature that would allow users to set a maximum token limit, but I convinced them to valide the need first.
User interviews and a smoke test, which displayed token usage after every response alongside a "We're working on it" message, revealed that users had little interest in monitoring token consumption. They wanted to know when they run out.
That allowed me to cut 20% of backlog saving our precious time.


Challenges
As a startup, we couldn't compete on model size or infrastructure. Instead, we focused on something Big Tech couldn't easily replicate: deep legal expertise with trustworthy sources.
Although Gaius-Lex offered dedicated workflows for legal professionals, a vast majority of users came looking for one thing: an AI assistant.
Building trust was challenged because of:
Unpredictable technology
The biggest challenge was fighting with models themselves. There were days were bugs surprised us each morning. A prompt that worked yesterday could fail a day after.
The expectation gap
Users expected the speed, quality, and pricing of products backed by multi-billion-dollar investments.
Everyone uses AI, no one understands it
People expected prompts like "Don't hallucinate the law" to guarantee accurate answers.
Others believed telling chat to "delete sensitive information” meant fulfilling their confidentiality obligations. Some expected the AI to "write more like a human" without providing examples or teaching the system what "human" meant for them.
Each ‘user error’ is a design opportunity.
Challenges
As a startup, we couldn't compete on model size or infrastructure. Instead, we focused on something Big Tech couldn't easily replicate: deep legal expertise with trustworthy sources.
Although Gaius-Lex offered dedicated workflows for legal professionals, a vast majority of users came looking for one thing: an AI assistant.
Building trust was challenged because of:
Unpredictable technology
The biggest challenge was fighting with models themselves. There were days were bugs surprised us each morning. A prompt that worked yesterday could fail a day after.
The expectation gap
Users expected the speed, quality, and pricing of products backed by multi-billion-dollar investments.
Everyone uses AI, no one understands it
People expected prompts like "Don't hallucinate the law" to guarantee accurate answers.
Others believed telling chat to "delete sensitive information” meant fulfilling their confidentiality obligations. Some expected the AI to "write more like a human" without providing examples or teaching the system what "human" meant for them.
Each ‘user error’ is a design opportunity.


learnings
80% of users aren't confident they're prompting correctly, being in the "ask, copy, paste" first stage of AI adoption.
Users do not know AI's limitations
It's possible to combine usability tests with user education/sales opportunity, but requires plan & great communication
learnings
80% of users aren't confident they're prompting correctly, being in the "ask, copy, paste" first stage of AI adoption.
Users do not know AI's limitations
It's possible to combine usability tests with user education/sales opportunity, but requires plan & great communication
learnings
80% of users aren't confident they're prompting correctly, being in the "ask, copy, paste" first stage of AI adoption.
Users do not know AI's limitations
It's possible to combine usability tests with user education/sales opportunity, but requires plan & great communication
Get in touch
Local time in Gdańsk, Poland
🇵🇱 🇪🇺
© 2025 Fembot Studio 👾 by Zuzanna Adamczyk. All rights reserved.
Get in touch
Local time in Gdańsk, Poland
🇵🇱 🇪🇺
© 2025 Fembot Studio 👾 by Zuzanna Adamczyk. All rights reserved.
Get in touch
Local time in Gdańsk, Poland
🇵🇱 🇪🇺
© 2025 Fembot Studio 👾 by Zuzanna Adamczyk. All rights reserved.