I built non-autoregressive decision models with RL a year ago
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
I built non-autoregressive decision models with RL a year ago
Unofficial Hacker News client; not affiliated with Y Combinator.
johnfn · · focus · HN ↗
OPs “marketing” is a single post on Reddit titled “ Predicting sales conversion probability from conversations using pure Reinforcement Learning”. Can you understand what that means? I can’t, and I consider myself reasonably technical. Is it obvious it has the same implications as Jev? Again, no idea. And it was just a single post on a subreddit that I don’t even browse! I see people on this thread saying “Jev is just BERT”. Sure, and Dropbox is just a ftp account mounted with curlftpfs!
I do feel bad for the author for finding something cool and being unable to brand it. But the full definition of “product” INCLUDES being able to coherently communicate it. In some sense the branding is just as much the “breakthrough” as the model.
rvz · · focus · HN ↗
No one cares if you are "first". They only care if your product is known by as many people as possible and is better than all the other alternatives at solving a problem that is worth paying for.
If you don't market, then no-one will care that you exist even if you solved a problem decades ago. Someone else will use your solution and take inspiration (and credit) off of your discovery because you didn't bother to tell anyone about it.
This is exactly what happened here.
verdverm · · focus · HN ↗
wavewrangler · · focus · HN ↗
malux85 · · focus · HN ↗
But that is NOT the point AT ALL. The point is, the same point that comes up on hacker news 1000 times a year - ideas alone are near worthless and execution matters.
Execution includes marketing that gets you enough attention. Because theres 10,000 other similar ideas of varying quality and marketing that others will point go saying "I wAS tHErE fiRsT"
All of this stems from the human bias of both (a) wishful thinking and (b)thinking people value what what we produce. These are natural human biases and are often dangerously wrong.
Programmers always think its just the idea and a prototype that is valuable, because they can produce and idea and a prototype (people what what I have) which causes them to massively overvalue ideas and the importance of "who was first" and all of that because they are sanctifying the small thing they produce.
It doesnt have to be malicious - the plain truth is theres 10,000 other ideas that are close enough that could be considered stealing even if they were truly independently developed, ideas are virtually worthless, get rid of your human biases that are clouding your judgment and focus on what matters if your idea is truly great : execution
wavewrangler · · focus · HN ↗
[dead]
bennett_dev · · focus · HN ↗
victor9000 · · focus · HN ↗
robrenaud · · focus · HN ↗
I can understand it, and it wouldn't excite me at all.
Jev has a beautiful API and is advertised as something much more general.
verdverm · · focus · HN ↗
(the project before it was rehashed into Laya since Jev was released)
tinyhouse · · focus · HN ↗
anonzzzies · · focus · HN ↗
verdverm · · focus · HN ↗
<a href="https://news.ycombinator.com/item?id=49674396">https://news.ycombinator.com/item?id=49674396
calebkaiser · · focus · HN ↗
Statistical modeling, from simple classical stuff up to modern deep learning, just has this dynamic where the theory is rich and bottomless, but the actual components of implementation are pretty neat and compact. So for any given idea, there are probably 20,000 other people who have had the same intuition, just with subtly different application or implementation. Add in that depending on what your particular flavor of research is, you might name an almost identical implementation something completely different. And it leads to a huge amount of sour grapes whenever anyone's idea really garners attention.
If you listen to any podcast with a founder in the ML space who has been in it for long enough, they will invariably say at some point "We actually developed xyz over a year before OpenAI"
hiddencost · · focus · HN ↗
TomGarden · · focus · HN ↗
So while the initial post was not good, the author is currently succeeding to some extent at what you're describing
XTXinverseXTY · · focus · HN ↗
It is arrogant and entitled for the author to take credit for the concept of RL over sequence embeddings, and none of the work that went into pretraining, not to mention the egregious target leakage [1]
[0]: Author fails to grasp the concept of virtual environments <a href="https://www.reddit.com/r/LocalLLaMA/comments/1kl0uvv/comment/ms0n31z/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button" rel="nofollow">https://www.reddit.com/r/LocalLLaMA/comments/1kl0uvv/comment...
[1]: his `train.py` has `outcome` as a model input (conversation_metrics built from _parse_conversation which includes outcome): <a href="https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning/blob/main/train.py#L73" rel="nofollow">https://huggingface.co/DeepMostInnovations/sales-conversion-... <a href="https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning/blob/main/train.py#L170" rel="nofollow">https://huggingface.co/DeepMostInnovations/sales-conversion-...
[2]: 100% of this post is AI-generated <a href="https://www.pangram.com/history/97e0be84-391d-46b8-9c16-2d8fcf8aebf5?ucc=X4GRCBL3tlK" rel="nofollow">https://www.pangram.com/history/97e0be84-391d-46b8-9c16-2d8f...
OceanKing · · focus · HN ↗
1)`outcome` is part of `metrics` at <a href="https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning/blob/main/train.py#L226" rel="nofollow">https://huggingface.co/DeepMostInnovations/sales-conversion-... and <a href="https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning/blob/main/train.py#L272" rel="nofollow">https://huggingface.co/DeepMostInnovations/sales-conversion-...
2) `metrics` goes into `ConversationState` at <a href="https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning/blob/main/train.py#L285" rel="nofollow">https://huggingface.co/DeepMostInnovations/sales-conversion-... and <a href="https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning/blob/main/train.py#L238" rel="nofollow">https://huggingface.co/DeepMostInnovations/sales-conversion-...
3) `metrics` (including `outcome`) makes its way into `ConversationState.state_vector` at <a href="https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning/blob/main/train.py#L73" rel="nofollow">https://huggingface.co/DeepMostInnovations/sales-conversion-..., and is returned from environment `step()` and `reset()` functions at <a href="https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning/blob/main/train.py#L243" rel="nofollow">https://huggingface.co/DeepMostInnovations/sales-conversion-... and <a href="https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning/blob/main/train.py#L290" rel="nofollow">https://huggingface.co/DeepMostInnovations/sales-conversion-...
4) model ingests `state_vector` as input at <a href="https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning/blob/main/train.py#L537" rel="nofollow">https://huggingface.co/DeepMostInnovations/sales-conversion-...
fxwin · · focus · HN ↗
I was curious about this so I skimmed the paper [0]:
> SalesRLAgent achieved 96.7% accuracy, outperforming the best commercial alternative by 23.7 percentage points and the best LLM approach by 34.7 percentage points.
For a fuzzy natural language task like this, this magnitude of improvement should already set off alarm bells (Though i admit I'm not even sure what accuracy is even measured here, and the paper doesn't help either). Also, "best LLM" here refers to GPT-4 (at the time of upload, the public already had access to GPT-o3 and). I would have loved to contextualize the performance by looking at model size, but the paper is frustratingly devoid of detail in that regard:
> The core of SalesRLAgent is a reinforcement learning architecture consisting of: • A state encoder network that processes Azure OpenAI embeddings and features • A policy network that estimates conversion probability based on the current state • A value network that estimates the expected cumulative reward • A meta-learning module that assesses prediction confi dence
Also:
> Beyond technical metrics, we evaluated SalesRLAgent in real-world sales environments through A/B testing. [...] After 90 days across 217 representatives and 12,433 con versations, we observed: • 43.2% increase in conversion rate for the test group
This would be a pretty huge result but the fact that this is just shoved into a single paragrpah with no further discussion on methodology, baselines and setup makes me very suspicious.
[0] <a href="https://arxiv.org/abs/2503.23303" rel="nofollow">https://arxiv.org/abs/2503.23303
porridgeraisin · · focus · HN ↗
Also, the way highly empirical fields like ML work is that it could very well be the case that typesafe had to do a _lot_ of work to improve this one, and in this field it ends up different enough that they feel they are doing something entirely novel[1]. I am not endorsing that 100%, but that happens a lot even between academics. In many cases it is valid.
[1] For example, this guys implementation seems to have atleast one serious issue, as {solution to OLS} points out in a sibling comment: <a href="https://news.ycombinator.com/item?id=49770027">https://news.ycombinator.com/item?id=49770027
tootie · · focus · HN ↗
I have no idea if I was literally the first person to have this idea or if anyone who launched one of these businesses read my article. I definitely didn't understand how much commerical value there was or even considered making a business out of it. I blame nobody but myself for missing an opportunity if there even was one.
I did get like $200 for writing it which was nice.
[deleted] · · focus · HN ↗
[deleted]
mcapodici · · focus · HN ↗
While they may not initially done well they are certainly riding this wave.
onion2k · · focus · HN ↗
HN readers are awesome at saying something is great and upvoting it, but unless HN readers are your market it means absolutely nothing. Marketing is as much about putting your message to the right audience as it is about saying the right thing.
This is made more complicated because a group as diverse as HN readers probably does contain some people who are in your target market, to be fair. The problem is that you're getting a strong signal from the whole cohort rather than the bit you're interested in, and it's really easy to conflate that with a sign of success.
As always with any startup activity, unless people are actually giving you their money it doesn't count and you should consider it a vanity metric.
derwiki · · focus · HN ↗
d--b · · focus · HN ↗
geuis · · focus · HN ↗
The usage of them immediately makes your commentary suspect. Either you aren't using the standard web interface to make a comment, you're using and odd 3rd party client, or its LLM generated.
isityettime · · focus · HN ↗
I don't give a shit if my Hacker News comments sometimes look "suspect" to some people. The way I use language is deeply personal and I'm not going to let the clankers or reactionaries against them take it away from me.
hbarka · · focus · HN ↗
swingboy · · focus · HN ↗
kwinkunks · · focus · HN ↗
johnfn · · focus · HN ↗
glerk · · focus · HN ↗
That's not just bad marketing, it's an example of anti-marketing.
Sales conversion? That makes me think of an old car's salesman trying to scam me into buying something I don't want. I positively don't want to read this paper based on the title.
WhitneyLand · · focus · HN ↗
This comparison doesn’t even make sense.
ranger_danger · · focus · HN ↗
emykhailenko · · focus · HN ↗